遇见数据集

daidalos-project/Herodotos_dataset

收藏
Hugging Face2024-10-06 更新2025-04-26 收录
官方服务:

资源简介:

--- license: agpl-3.0 --- Datasheet: Herodotos Project Dataset For what purpose was the dataset created? Was there a specific task in mind? Was there a specific gap that needed to be filled? Please provide a description. - created for Herodotos Project to train NER-Tagger (BiLSTM CRF; see: Alexander Erdmann, David Joseph Wrisley, Benjamin Allen, Christopher Brown, Sophie Cohen Bodénès, Micha Elsner, Yukun Feng, Brian Joseph, Béatrice Joyeaux-Prunel and Marie-Catherine de Marneffe. 2019. "[Practical, Efficient, and Customizable Active Learning for Named Entity Recognition in the Digital Humanities](https://github.com/alexerdmann/HER/blob/master/HER_NAACL2019_preprint.pdf)." In Proceedings of North American Association of Computational Linguistics (NAACL 2019). Minneapolis, Minnesota.) - Goal of Herodotos Project: catalogue and compendium of ancient ethnic groups - For more info on the corpus see: https://aclanthology.org/W16-4012.pdf Who created the dataset (e.g., which team, research group) and on behalf of which entity (e.g., company, institution, organization)? - from the documentation: „The data files in the **Annotation** directory were annotated for named entities by a team of Classics experts at Ohio State University. Texts presently included are excerpts from Caesar's Wars, both Gallic (GW) and Civil (CW), the Plinies' writings, both Elder and Younger, and Ovid's Ars Amatoria. " Who funded the creation of the dataset? If there is an associated grant, please provide the name of the grantor and the grant name and number. - unknown Any other comments? - No What do the instances that comprise the dataset represent (e.g., documents, photos, people, countries)? Are there multiple types of instances (e.g., movies, users, and ratings; people and interactions between them; nodes and edges)? Please provide a description. - Latin texts "Texts presently included are excerpts from Caesar's Wars, both Gallic (GW) and Civil (CW), the Plinies' writings, both Elder and Younger, and Ovid's Ars Amatoria." How many instances are there in total (of each type, if appropriate)? - 146,066 words Does the dataset contain all po-ssible instances or is it a sample (not necessarily random) of instances from a larger set? If the dataset is a sample, then what is the larger set? Is the sample representative of the larger set (e.g., geographic coverage)? If so, please describe how this representativeness was validated/verified. If it is not representative of the larger set, please describe why not (e.g., to cover a more diverse range of instances, because instances were withheld or unavailable). - sample of Latin literature (see previous answers), representative of Classical Latin literature, might not be representative of the entire Latin literature (time, geography) What data does each instance consist of? "Raw" data (e.g., unprocessed text or images) or features? In either case, please provide a description. - Each instance consists of raw text data Is there a label or target associated with each instance? If so, please provide a description. - NER Labels: PRS-B, PRS-I, GEO-B, GEO-I, GRP-B, GRP-I or 0 - labels follow the BIO scheme - see also: https://aclanthology.org/W16-4012.pdf Is any information missing from individual instances? If so, please provide a description, explaining why this information is missing (e.g., because it was unavailable). This does not include intentionally removed information, but might include, e.g., redacted text. - No Are relationships between individual instances made explicit (e.g., users' movie ratings, social network links)? If so, please describe how these relationships are made explicit. - Relationships are made explicit according to the BIO scheme Are there recommended data splits (e.g., training, development/validation, testing)? If so, please provide a description of these splits, explaining the rationale behind them. - Text from Gallic War is split into test and train sets Are there any errors, sources of noise, or redundancies in the dataset? If so, please provide a description. - Naturally occurring repetitions of names in the texts Is the dataset self-contained, or does it link to or otherwise rely on external resources (e.g., websites, tweets, other datasets)? If it links to or relies on external resources, a) are there guarantees that they will exist, and remain constant, over time; b) are there official archival versions of the complete dataset (i.e., including the external resources as they existed at the time the dataset was created); c) are there any restrictions (e.g., licenses, fees) associated with any of the external resources that might apply to a dataset consumer? Please provide descriptions of all external resources and any restrictions associated with them, as well as links or other access points, as appropriate. - The dataset is self-contained and can be downloaded from GitHub (https://github.com/Herodotos-Project/Herodotos-Project-Latin-NER-Tagger-Annotation/blob/master/README.md) Does the dataset contain data that might be considered confidential (e.g., data that is protected by legal privilege or by doctor--patient confidentiality, data that includes the content of individuals' non-public communications)? If so, please provide a description. - No Does the dataset contain data that, if viewed directly, might be offensive, insulting, threatening, or might otherwise cause anxiety? If so, please describe why. If the dataset does not relate to people, you may skip the remaining questions in this section. - The dataset contains descriptions of war. Does the dataset identify any subpopulations (e.g., by age, gender)? If so, please describe how these subpopulations are identified and provide a description of their respective distributions within the dataset. - A number of ethnic groups from antiquity are referred to. Is it possible to identify individuals (i.e., one or more natural persons), either directly or indirectly (i.e., in combination with other data) from the dataset? If so, please describe how. - Only historical individuals Does the dataset contain data that might be considered sensitive in any way (e.g., data that reveals race or ethnic origins, sexual orientations, religious beliefs, political opinions or union memberships, or locations; financial or health data; biometric or genetic data; forms of government identification, such as social security numbers; criminal history)? If so, please provide a description. - Only historical individuals Any other comments? - No How was the data associated with each instance acquired? Was the data directly observable (e.g., raw text, movie ratings), reported by subjects (e.g., survey responses), or indirectly inferred/derived from other data (e.g., part-of-speech tags, model-based guesses for age or language)? - The data consists of publicly available texts If the data was reported by subjects or indirectly inferred/derived from other data, was the data validated/verified? If so, please describe how. - unknown What mechanisms or procedures were used to collect the data (e.g., hardware apparatuses or sensors, manual human curation, software programs, software APIs)? How were these mechanisms or procedures validated? - from the documentation: „All texts are in Latin taken from the [Latin Library Collection](https://www.theLatinlibrary.com/) (collected by [CLTK](https://github.com/cltk/Latin_text_Latin_library)) or the [Perseus Latin Collection](http://www.perseus.tufts.edu/hopper/collection?collection=Perseus:collection:Greco-Roman). " If the dataset is a sample from a larger set, what was the sampling strategy (e.g., deterministic, probabilistic with specific sampling probabilities)? - unknown Who was involved in the data collection process (e.g., students, crowdworkers, contractors) and how were they compensated (e.g., how much were crowdworkers paid)? - <https://aclanthology.org/W16-4012.pdf> S. 87: "an undergraduate, a graduate, and a professor of Classics, each with at least 4 years of experience studying Latin" Over what timeframe was the data collected? Does this timeframe match the creation timeframe of the data associated with the instances (e.g., recent crawl of old news articles)? If not, please describe the timeframe in which the data associated with the instances was created. Were any ethical review processes conducted (e.g., by an institutional review board)? If so, please provide a description of these review processes, including the outcomes, as well as a link or other access point to any supporting documentation. If the dataset does not relate to people, you may skip the remaining questions in this section. - unknown Did you collect the data from the individuals in question directly, or obtain it via third parties or other sources (e.g., websites)? Were the individuals in question notified about the data collection? If so, please describe (or show with screenshots or other information) how notice was provided, and provide a link or other access point to, or otherwise reproduce, the exact language of the notification itself. - not applicable Did the individuals in question consent to the collection and use of their data? If so, please describe (or show with screenshots or other information) how consent was requested and provided, and provide a link or other access point to, or otherwise reproduce, the exact language to which the individuals consented. - not applicable If consent was obtained, were the consenting individuals provided with a mechanism to revoke their consent in the future or for certain uses? If so, please provide a description, as well as a link or other access point to the mechanism (if appropriate). - not applicable Has an analysis of the potential impact of the dataset and its use on data subjects (e.g., a data protection impact analysis) been conducted? If so, please provide a description of this analysis, including the outcomes, as well as a link or other access point to any supporting documentation. - not applicable Any other comments? - No Preprocessing/cleaning/labeling Was any preprocessing/cleaning/labeling of the data done (e.g., discretization or bucketing, tokenization, part-of-speech tagging, SIFT feature extraction, removal of instances, processing of missing values)? If so, please provide a description. If not, you may skip the remainder of the questions in this section. - The data was manually annotated for NEs. Was the "raw" data saved in addition to the preprocessed/cleaned/labeled data (e.g., to support unanticipated future uses)? If so, please provide a link or other access point to the "raw" data. - The data can be downloaded from: https://github.com/clmarr/Herodotos-beta/tree/f22fdd92b3318cfb8fc93b004b0947aea14ce9c2/Annotation_1-1-19 Any other comments? -No Uses Has the dataset been used for any tasks already? If so, please provide a description. - It has been used to train an NER-Tagger for Latin. See: <https://aclanthology.org/W16-4012.pdf> and https://github.com/alexerdmann/HER/blob/master/HER_NAACL2019_preprint.pdf Is there a repository that links to any or all papers or systems that use the dataset? If so, please provide a link or other access point. What (other) tasks could the dataset be used for? - See: https://github.com/alexerdmann/HER/blob/master/HER_NAACL2019_preprint.pdf Is there anything about the composition of the dataset or the way it was collected and preprocessed/cleaned/labeled that might impact future uses? For example, is there anything that a dataset consumer might need to know to avoid uses that could result in unfair treatment of individuals or groups (e.g., stereotyping, quality of service issues) or other risks or harms (e.g., legal risks, financial harms)? If so, please provide a description. Is there anything a dataset consumer could do to mitigate these risks or harms? - Strong class imbalance (most tokens are non-entities) Are there tasks for which the dataset should not be used? If so, please provide a description. - No Any other comments? - No Distribution Will the dataset be distributed to third parties outside of the entity (e.g., company, institution, organization) on behalf of which the dataset was created? If so, please provide a description. How will the dataset will be distributed (e.g., tarball on website, API, GitHub)? - The data can be downloaded from: https://github.com/clmarr/Herodotos-beta/tree/f22fdd92b3318cfb8fc93b004b0947aea14ce9c2/Annotation_1-1-19 Does the dataset have a digital object identifier (DOI)? - No When will the dataset be distributed? - The data can be downloaded from: - <https://github.com/clmarr/Herodotos-beta/tree/f22fdd92b3318cfb8fc93b004b0947aea14ce9c2/Annotation_1-1-19> - <https://github.com/Herodotos-Project/Herodotos-Project-Latin-NER-Tagger-Annotation> Will the dataset be distributed under a copyright or other intellectual property (IP) license, and/or under applicable terms of use (ToU)? If so, please describe this license and/or ToU, and provide a link or other access point to, or otherwise reproduce, any relevant licensing terms or ToU, as well as any fees associated with these restrictions. - [AGPL-3.0 license](https://github.com/Herodotos-Project/Herodotos-Project-Latin-NER-Tagger-Annotation/blob/master/LICENSE) Have any third parties imposed IP-based or other restrictions on the data associated with the instances? If so, please describe these restrictions, and provide a link or other access point to, or otherwise reproduce, any relevant licensing terms, as well as any fees associated with these restrictions. - unknown Do any export controls or other regulatory restrictions apply to the dataset or to individual instances? If so, please describe these restrictions, and provide a link or other access point to, or otherwise reproduce, any supporting documentation. - unknown Any other comments? - No Maintenance Who will be supporting/hosting/maintaining the dataset? - from the documentation: "Contact [ae1541@nyu.edu](mailto:ae1541@nyu.edu) or any of the co-authors with questions regarding this repository." How can the owner/curator/manager of the dataset be contacted (e.g., email address)? - [ae1541@nyu.edu](mailto:ae1541@nyu.edu) Is there an erratum? If so, please provide a link or other access point. Will the dataset be updated (e.g., to correct labeling errors, add new instances, delete instances)? If so, please describe how often, by whom, and how updates will be communicated to dataset consumers (e.g., mailing list, GitHub)? - new instances for the Ancient Greek language will be added in the future If the dataset relates to people, are there applicable limits on the retention of the data associated with the instances (e.g., were the individuals in question told that their data would be retained for a fixed period of time and then deleted)? If so, please describe these limits and explain how they will be enforced. - not applicable Will older versions of the dataset continue to be supported/hosted/maintained? If so, please describe how. If not, please describe how its obsolescence will be communicated to dataset consumers. - unknown If others want to extend/augment/build on/contribute to the dataset, is there a mechanism for them to do so? If so, please provide a description. Will these contributions be validated/verified? If so, please describe how. If not, why not? Is there a process for communicating/distributing these contributions to dataset consumers? If so, please provide a description. - unknown Any other comments? - No

数据集说明:希罗多德项目数据集(Herodotos Project Dataset) --- 许可证:AGPL-3.0 --- ### 数据集创建目的 本数据集创建的目的是什么?是否有特定的目标任务?是否存在需要填补的特定研究空白?请予以说明。 - 本数据集为希罗多德项目创建,用于训练命名实体识别标注器(NER-Tagger,采用双向长短期记忆网络条件随机场(BiLSTM CRF)架构;参见:Alexander Erdmann、David Joseph Wrisley、Benjamin Allen、Christopher Brown、Sophie Cohen Bodénès、Micha Elsner、Yukun Feng、Brian Joseph、Béatrice Joyeaux-Prunel 及 Marie-Catherine de Marneffe. 2019. 《[实用、高效且可定制的主动学习方法,用于数字人文领域的命名实体识别](https://github.com/alexerdmann/HER/blob/master/HER_NAACL2019_preprint.pdf)》,载于《北美计算语言学协会会议录(NAACL 2019)》,明尼阿波利斯,明尼苏达州。) - 希罗多德项目的目标:构建古代族群的目录与汇编 - 如需了解语料库更多信息,请访问:https://aclanthology.org/W16-4012.pdf ### 数据集创建主体 本数据集由谁创建(例如所属团队、研究小组),代表哪个实体(例如公司、机构、组织)? - 根据文档说明:**标注(Annotation)** 目录下的数据文件由俄亥俄州立大学古典学专家团队完成命名实体标注。当前收录的文本包括凯撒《战记》(高卢战记(GW)与内战记(CW))、老普林尼与小普林尼的著作,以及奥维德的《爱经》(*Ars Amatoria*)。 ### 资助情况 本数据集的创建是否有资金支持?若有相关资助项目,请提供资助方名称、资助项目名称及编号。 - 未知 ### 其他说明 无 ### 实例含义与规模 本数据集包含的实例分别代表什么(例如文档、图片、人物、国家)?是否存在多种类型的实例(例如电影、用户与评分;人物及其交互关系;节点与边)?请予以说明。 - 拉丁语文本:当前收录的文本包括凯撒《战记》(高卢战记(GW)与内战记(CW))、老普林尼与小普林尼的著作,以及奥维德的《爱经》。 总实例数量(若适用,分类型说明): - 共146,066个词 数据集是否包含全部可能的实例,还是仅为更大规模语料库中的采样样本(未必是随机采样)?若为采样样本,那么更大规模的语料库是什么?该样本是否能代表更大规模的语料库(例如地理覆盖范围)?若是,请描述该代表性的验证/确认方式;若否,请说明原因(例如为覆盖更多样化的实例、因部分实例被保留或无法获取)。 - 本数据集为拉丁文学术语料的采样样本(详见前文说明),可代表古典拉丁语文学,但无法代表全部拉丁语文学(在时间、地理维度上存在局限)。 ### 实例数据构成 每个实例包含何种数据?是原始数据(例如未处理的文本或图片)还是特征?无论属于哪种情况,请予以说明。 - 每个实例均包含原始文本数据。 每个实例是否关联有标签或目标变量?若是,请予以说明。 - 命名实体识别标签:PRS-B、PRS-I、GEO-B、GEO-I、GRP-B、GRP-I 或 0 - 标签遵循BIO标注体系 - 亦可参考:https://aclanthology.org/W16-4012.pdf 单个实例是否存在信息缺失?若是,请说明原因(例如因无法获取)。此处不包括有意移除的信息,但可能包括例如打码文本等情况。 - 无 实例间的关系是否已明确(例如用户的电影评分、社交网络链接)?若是,请描述这些关系的明确方式。 - 实例间的关系已通过BIO标注体系明确。 是否有推荐的数据划分方式(例如训练集、开发/验证集、测试集)?若是,请描述这些划分方式,并说明其背后的逻辑依据。 - 高卢战记的文本被划分为测试集与训练集。 数据集是否存在错误、噪声源或冗余内容?若是,请予以说明。 - 文本中存在自然出现的名称重复现象。 ### 外部资源依赖情况 数据集是自包含的,还是会链接至或依赖外部资源(例如网站、推文、其他数据集)?若存在链接或依赖外部资源的情况:a) 是否能保证这些资源长期存在且稳定;b) 是否存在完整数据集的官方存档版本(即包含数据集创建时的外部资源);c) 数据集消费者是否需要遵守与任何外部资源相关的限制(例如许可证、费用)?请描述所有外部资源及相关限制,以及适用的链接或其他访问途径。 - 本数据集为自包含格式,可从GitHub下载:https://github.com/Herodotos-Project/Herodotos-Project-Latin-NER-Tagger-Annotation/blob/master/README.md ### 敏感数据相关 数据集是否包含可能被视为机密的数据(例如受法律特权保护的数据、医患保密数据、包含个人非公开通信内容的数据)?若是,请予以说明。 - 无 数据集包含的内容若直接查看,是否可能具有冒犯性、侮辱性、威胁性,或引发其他焦虑情绪?若是,请说明原因。若数据集与人物无关,可跳过本小节剩余问题。 - 本数据集包含战争相关描述。 数据集是否识别出任何子群体(例如按年龄、性别划分)?若是,请描述这些子群体的识别方式,并说明其在数据集中的分布情况。 - 本数据集提及了多个古代族群。 是否可以直接或间接(例如结合其他数据)从数据集中识别出个体(即一个或多个自然人)?若是,请说明识别方式。 - 仅涉及历史人物。 数据集是否包含任何可能被视为敏感的数据(例如揭示种族或族裔起源、性取向、宗教信仰、政治观点或工会成员身份的数据;财务或健康数据;生物特征或遗传数据;政府身份标识,例如社会保障号码;犯罪记录)?若是,请予以说明。 - 仅涉及历史人物。 其他说明:无 ### 数据获取与收集流程 每个实例关联的数据是如何获取的?数据是直接可观测的(例如原始文本、电影评分)、由主体报告的(例如调查回复),还是从其他数据间接推断/衍生得到的(例如词性标注、基于模型的年龄或语言猜测)? - 本数据集的数据均来自公开可用的文本。 若数据由主体报告或从其他数据间接推断/衍生得到,是否对数据进行过验证?若是,请说明验证方式。 - 未知 采用何种机制或流程收集数据(例如硬件设备或传感器、人工整理、软件程序、软件API)?这些机制或流程是否经过验证?如何验证的? - 根据文档说明:所有文本均取自[拉丁语馆藏库](https://www.theLatinlibrary.com/)(由古典语言工具包(CLTK)整理)或[珀尔修斯拉丁语馆藏库](http://www.perseus.tufts.edu/hopper/collection?collection=Perseus:collection:Greco-Roman)。 若数据集为更大规模语料库的采样样本,那么采样策略是什么(例如确定性采样、带有特定采样概率的概率采样)? - 未知 参与数据收集流程的人员有哪些(例如学生、众包工作者、承包商)?他们的报酬方式是什么(例如众包工作者的薪酬标准)? - 引自https://aclanthology.org/W16-4012.pdf 第87页:一名本科生、一名研究生及一名古典学教授,每位均拥有至少4年的拉丁语学习研究经验 数据收集的时间范围是什么?该时间范围是否与实例关联数据的创建时间范围一致(例如近期爬取的旧新闻文章)?若不一致,请说明实例关联数据的创建时间范围。是否进行过任何伦理审查流程(例如由机构审查委员会进行)?若是,请描述这些审查流程及结果,并提供任何支持文档的链接或其他访问途径。若数据集与人物无关,可跳过本小节剩余问题。 - 未知 您是直接从相关个体收集数据,还是通过第三方或其他来源(例如网站)获取数据?相关个体是否被告知数据收集事宜?若是,请描述(或附上截图或其他信息)通知的方式,并提供通知原文的链接或其他访问途径,或直接复现通知的准确措辞。 - 不适用 相关个体是否同意收集和使用其数据?若是,请描述(或附上截图或其他信息)请求和获得同意的方式,并提供同意原文的链接或其他访问途径,或直接复现同意的准确措辞。 - 不适用 若已获得同意,是否向同意的个体提供了未来撤销同意或针对特定使用场景撤销同意的机制?若是,请描述该机制,并提供相关链接或其他访问途径(如适用)。 - 不适用 是否对数据集及其使用对数据主体的潜在影响进行过分析(例如数据保护影响分析)?若是,请描述该分析及结果,并提供任何支持文档的链接或其他访问途径。 - 不适用 其他说明:无 ### 预处理/清洗/标注 是否对数据进行过任何预处理、清洗或标注操作(例如离散化或分箱、分词、词性标注、SIFT特征提取、移除实例、处理缺失值)?若是,请予以说明。若未进行,请跳过本小节剩余问题。 - 本数据集已针对命名实体完成人工标注。 是否除了预处理/清洗/标注后的数据之外,还保存了原始数据(例如为支持未来未预期的使用场景)?若是,请提供原始数据的链接或其他访问途径。 - 本数据集可从以下地址下载:https://github.com/clmarr/Herodotos-beta/tree/f22fdd92b3318cfb8fc93b004b0947aea14ce9c2/Annotation_1-1-19 其他说明:无 ### 数据集使用情况 本数据集是否已被用于任何任务?若是,请予以说明。 - 本数据集已被用于训练拉丁语命名实体识别标注器。详见:https://aclanthology.org/W16-4012.pdf 及 https://github.com/alexerdmann/HER/blob/master/HER_NAACL2019_preprint.pdf 是否存在与使用本数据集的所有论文或系统相关的代码仓库?若是,请提供链接或其他访问途径。该数据集还可用于哪些(其他)任务? - 详见:https://github.com/alexerdmann/HER/blob/master/HER_NAACL2019_preprint.pdf 数据集的组成、收集方式或预处理/清洗/标注流程是否存在可能影响未来使用的问题?例如,数据集消费者是否需要了解某些信息以避免可能导致个体或群体受到不公平对待的使用场景(例如刻板印象、服务质量问题)或其他风险或危害(例如法律风险、财务损失)?若是,请予以说明。数据集消费者可采取何种措施来减轻这些风险或危害? - 存在严重的类别不平衡问题(绝大多数词元均为非实体)。 是否存在本数据集不应被用于的任务?若是,请予以说明。 - 无 其他说明:无 ### 数据集分发 本数据集是否会分发给创建该数据集的实体之外的第三方(例如公司、机构、组织)?若是,请予以说明。数据集将以何种方式分发(例如网站上的tar包、API、GitHub)? - 本数据集可从以下地址下载:https://github.com/clmarr/Herodotos-beta/tree/f22fdd92b3318cfb8fc93b004b0947aea14ce9c2/Annotation_1-1-19 本数据集是否拥有数字对象标识符(DOI)? - 无 本数据集何时会被分发? - 本数据集可从以下地址下载: - https://github.com/clmarr/Herodotos-beta/tree/f22fdd92b3318cfb8fc93b004b0947aea14ce9c2/Annotation_1-1-19 - https://github.com/Herodotos-Project/Herodotos-Project-Latin-NER-Tagger-Annotation 本数据集是否会在版权或其他知识产权(IP)许可下分发,以及/或适用相关使用条款(ToU)?若是,请描述该许可证和/或使用条款,并提供相关许可条款或使用条款的链接或其他访问途径,或直接复现相关内容,以及与这些限制相关的任何费用。 - [AGPL-3.0许可证](https://github.com/Herodotos-Project/Herodotos-Project-Latin-NER-Tagger-Annotation/blob/master/LICENSE) 是否有第三方对实例关联的数据施加了基于IP的或其他限制?若是,请描述这些限制,并提供相关许可条款的链接或其他访问途径,或直接复现相关内容,以及与这些限制相关的任何费用。 - 未知 是否有出口管制或其他监管限制适用于本数据集或单个实例?若是,请描述这些限制,并提供支持文档的链接或其他访问途径,或直接复现相关内容。 - 未知 其他说明:无 ### 数据集维护 谁将支持、托管或维护本数据集? - 根据文档说明:如有关于本仓库的疑问,请联系[ae1541@nyu.edu](mailto:ae1541@nyu.edu)或任何一位合著作者。 如何联系数据集的所有者、管理者或维护者(例如电子邮件地址)? - [ae1541@nyu.edu](mailto:ae1541@nyu.edu) 是否存在勘误表?若是,请提供链接或其他访问途径。本数据集是否会进行更新(例如修正标注错误、添加新实例、删除实例)?若是,请描述更新的频率、更新者,以及将如何向数据集消费者传达更新信息(例如邮件列表、GitHub)。 - 未来将添加古希腊语相关的新实例。 若数据集与人物相关,是否有适用于实例关联数据留存的适用限制(例如相关个体被告知其数据将留存固定时长后删除)?若是,请描述这些限制并说明将如何执行。 - 不适用 旧版本的数据集是否会继续得到支持、托管或维护?若是,请描述具体方式。若否,请描述将如何向数据集消费者传达其过时的信息。 - 未知 若其他方希望扩展、增强、基于本数据集进行开发或为本数据集做出贡献,是否存在相应的机制?若是,请予以描述。这些贡献是否会经过验证/确认?若是,请描述验证方式。若否,请说明原因。是否存在向数据集消费者传达/分发这些贡献的流程?若是,请予以描述。 - 未知 其他说明:无

提供机构:
daidalos-project
二维码
社区交流群
二维码
科研交流群
商业服务