Multimodal-Fatima/VizWiz_train
收藏资源简介:
--- dataset_info: features: - name: id dtype: int32 - name: image dtype: image - name: filename dtype: string - name: question dtype: string - name: answers sequence: string - name: answers_original list: - name: answer dtype: string - name: answer_confidence dtype: string - name: answer_type dtype: string - name: answerable dtype: int32 - name: id_image dtype: int64 - name: clip_tags_ViT_L_14 sequence: string - name: clip_tags_LAION_ViT_H_14_2B sequence: string - name: blip_caption_beam_5 dtype: string - name: LLM_Description_gpt3_downstream_tasks_visual_genome_ViT_L_14 sequence: string - name: LLM_Description_gpt3_downstream_tasks_visual_genome_LAION-ViT-H-14-2B sequence: string - name: DETA_detections_deta_swin_large_o365_coco_classes list: - name: attribute dtype: string - name: box sequence: float32 - name: label dtype: string - name: location dtype: string - name: ratio dtype: float32 - name: size dtype: string - name: tag dtype: string splits: - name: train num_bytes: 9906518637.0 num_examples: 20523 download_size: 9880125036 dataset_size: 9906518637.0 --- # Dataset Card for "VizWiz_train" [More Information needed](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
数据集信息: 特征列表: - 特征名:id,数据类型:int32 - 特征名:image,数据类型:图像 - 特征名:filename,数据类型:字符串 - 特征名:question,数据类型:字符串 - 特征名:answers,数据类型:字符串序列 - 特征名:answers_original,为嵌套列表结构,包含子特征: - 特征名:answer,数据类型:字符串 - 特征名:answer_confidence,数据类型:字符串 - 特征名:answer_type,数据类型:字符串 - 特征名:answerable,数据类型:int32 - 特征名:id_image,数据类型:int64 - 特征名:clip_tags_ViT_L_14,数据类型:CLIP标签(ViT_L_14)字符串序列 - 特征名:clip_tags_LAION_ViT_H_14_2B,数据类型:LAION-CLIP标签(ViT_H_14_2B)字符串序列 - 特征名:blip_caption_beam_5,数据类型:字符串 - 特征名:LLM_Description_gpt3_downstream_tasks_visual_genome_ViT_L_14,数据类型:大语言模型(Large Language Model,LLM)描述序列,基于GPT-3、下游任务为视觉基因组(Visual Genome)与ViT_L_14模型 - 特征名:LLM_Description_gpt3_downstream_tasks_visual_genome_LAION-ViT-H-14-2B,数据类型:大语言模型描述序列,基于GPT-3、下游任务为视觉基因组(Visual Genome)与LAION-ViT-H-14-2B模型 - 特征名:DETA_detections_deta_swin_large_o365_coco_classes,为嵌套列表结构,包含子特征: - 特征名:attribute,数据类型:字符串 - 特征名:box,数据类型:float32序列 - 特征名:label,数据类型:字符串 - 特征名:location,数据类型:字符串 - 特征名:ratio,数据类型:float32 - 特征名:size,数据类型:字符串 - 特征名:tag,数据类型:字符串 数据集划分: - 划分名称:train,字节占用量:9906518637.0,样本数量:20523 下载总大小:9880125036 数据集总占用大小:9906518637.0 --- # 数据集卡片:"VizWiz_train" [需补充更多信息](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
数据集概述
数据集名称
- VizWiz_train
数据集特征
- id:整数类型 (int32)
- image:图像类型
- filename:字符串类型 (string)
- question:字符串类型 (string)
- answers:字符串序列
- answers_original:列表类型,包含:
- answer:字符串类型 (string)
- answer_confidence:字符串类型 (string)
- answer_type:字符串类型 (string)
- answerable:整数类型 (int32)
- id_image:长整数类型 (int64)
- clip_tags_ViT_L_14:字符串序列
- clip_tags_LAION_ViT_H_14_2B:字符串序列
- blip_caption_beam_5:字符串类型 (string)
- LLM_Description_gpt3_downstream_tasks_visual_genome_ViT_L_14:字符串序列
- LLM_Description_gpt3_downstream_tasks_visual_genome_LAION-ViT-H-14-2B:字符串序列
- DETA_detections_deta_swin_large_o365_coco_classes:列表类型,包含:
- attribute:字符串类型 (string)
- box:浮点数序列 (float32)
- label:字符串类型 (string)
- location:字符串类型 (string)
- ratio:浮点数类型 (float32)
- size:字符串类型 (string)
- tag:字符串类型 (string)
数据集分割
- train:
- 数据量:9906518637.0 字节
- 示例数量:20523
数据集大小
- 下载大小:9880125036 字节
- 数据集大小:9906518637.0 字节




