napsternxg/nyt_ingredients
收藏资源简介:
New York Times Ingredient Phrase Tagger Dataset是一个用于从非结构化成分短语中提取数量、单位、名称和评论的数据集。数据集由专家生成,语言为英语,创建者未明确说明,但语言是从现有资源中找到的。数据集是单语言的,遵循Apache 2.0许可证。数据集的大小在10万到100万之间,标签包括食谱和成分,任务类别为令牌分类,具体任务为命名实体识别。数据集的原始来源是纽约时报的一个GitHub仓库,该仓库使用条件随机场模型(CRF)从标记的训练数据中提取标签,这些数据由人类新闻助理标记。
The New York Times Ingredient Phrase Tagger Dataset is a dataset designed for extracting quantities, units, ingredient names and comments from unstructured ingredient phrases. The dataset was generated by experts, is in English, and its creator is not explicitly specified, while the language materials were sourced from existing resources. It is a monolingual dataset released under the Apache 2.0 license. The dataset has a size ranging from 100,000 to 1,000,000 samples. Its labels cover recipes and ingredients, and it falls under the task category of token classification, specifically named entity recognition (NER). The original source of the dataset is a GitHub repository maintained by The New York Times. This repository uses the Conditional Random Fields (CRF) model to extract labels from annotated training data, which were labeled by human news assistants.
数据集概述
基本信息
- 名称: New York Times Ingredient Phrase Tagger Dataset
- 语言: 英语 (en)
- 语言创建者: 发现 (found)
- 许可证: Apache-2.0
- 多语言性: 单语 (monolingual)
- 大小: 10万<n<100万
详细描述
- 标签:
- 食谱 (recipe)
- 成分 (ingredients)
- 任务类别:
- 令牌分类 (token-classification)
- 任务ID:
- 命名实体识别 (named-entity-recognition)
数据来源
数据集创建
- 注释创建者: 专家生成 (expert-generated)
- 数据处理方法: 使用条件随机场模型 (CRF) 从标记的训练数据中提取标签,该数据由人类新闻助理标记。
数据集用途
- 用于从非结构化的成分短语中提取数量、单位、名称和评论,并应用于烹饪以格式化传入的食谱。



