MMID(Massively Multilingual Image Dataset)
收藏资源简介:
MMID 是一个大规模、大规模的多语言图像数据集,与在宾夕法尼亚大学收集的它们所代表的单词配对。数据集是双重并行的:对于每种语言,单词与表示单词的图像并行存储,并且与单词翻译成英语(和相应的图像)并行存储。 迄今为止最大的同类数据集,它有 98 种语言(包括英语),每种语言多达 10,000 个单词! (还有更多的英语。)
MMID is a large-scale multilingual image dataset, with each image paired to the word it represents, and the entire dataset was collected at the University of Pennsylvania. The dataset has a dual parallel structure: for each language, the source word is stored in parallel with the images that represent it, and additionally in parallel with the English translation of that word along with its corresponding images. As the largest dataset of its kind to date, it includes 98 languages (including English), with up to 10,000 words per language! (There are even more word entries for English.)




