Webis Cross-Lingual Sentiment Dataset 2010 (Webis-CLS-10)
收藏资源简介:
The Cross-Lingual Sentiment (CLS) dataset comprises about 800.000 Amazon product reviews in the four languages English, German, French, and Japanese. For more information on the construction of the dataset see (Prettenhofer and Stein, 2010) or the enclosed readme files. If you have a question after reading the paper and the readme files, please contact Peter Prettenhofer. We provide the dataset in two formats: 1) a processed format which corresponds to the preprocessing (tokenization, etc.) in (Prettenhofer and Stein, 2010); 2) an unprocessed format which contains the full text of the reviews (e.g., for machine translation or feature engineering). The dataset was first used by (Prettenhofer and Stein, 2010). It consists of Amazon product reviews for three product categories---books, dvds and music---written in four different languages: English, German, French, and Japanese. The German, French, and Japanese reviews were crawled from Amazon in November, 2009. The English reviews were sampled from the Multi-Domain Sentiment Dataset (Blitzer et. al., 2007). For each language-category pair there exist three sets of training documents, test documents, and unlabeled documents. The training and test sets comprise 2.000 documents each, whereas the number of unlabeled documents varies from 9.000 - 170.000.
跨语言情感(Cross-Lingual Sentiment, CLS)数据集包含英语、德语、法语及日语四种语言的约80万条亚马逊商品评论。如需了解数据集构建的更多细节,可参阅(Prettenhofer与Stein, 2010)或随附的自述文件。若在阅读论文与自述文件后仍有疑问,请联系Peter Prettenhofer。 本数据集提供两种存储格式:1)预处理格式,该格式对应(Prettenhofer与Stein, 2010)中提及的预处理流程(含分词等操作);2)未处理格式,包含评论的完整原文,可用于机器翻译或特征工程。 该数据集首次由(Prettenhofer与Stein, 2010)使用,涵盖图书、DVD及音乐三大商品品类的亚马逊评论,语种覆盖英语、德语、法语与日语四种语言。 其中德语、法语与日语评论爬取自2009年11月的亚马逊平台,英语评论则取自多领域情感数据集(Multi-Domain Sentiment Dataset, Blitzer et al., 2007)。 针对每一种语言-品类组合,均设有训练文档集、测试文档集与无标注文档集三个子集。训练集与测试集各包含2000条文档,而无标注集的文档数量则在9000至170000条之间浮动。



