ParlaCAP
收藏资源简介:
ParlaCAP是一个由约瑟夫·斯蒂芬研究所等机构联合创建的大规模多语言议会数据集,旨在分析欧洲28个国家和自治地区的议会议程设置。该数据集基于ParlaMint语料库,包含超过800万条议会演讲记录,涵盖20多种语言,数据量达12亿词。数据集创建过程采用了教师-学生框架,利用大型语言模型(如GPT-4o)自动标注政策主题标签,并通过多语言编码器模型进行微调以提升标注效率。ParlaCAP不仅提供政策主题分类,还包含丰富的演讲者及政党元数据,并整合了ParlaSent多语言情感分析模型的预测结果,为比较政治学研究提供了全面资源,可用于分析政治注意力分配、情感模式及政策关注中的性别差异等问题。
ParlaCAP is a large-scale multilingual parliamentary dataset jointly created by the Jožef Stefan Institute and other institutions, aiming to analyze parliamentary agenda-setting across 28 European countries and autonomous regions. Based on the ParlaMint corpus, this dataset contains over 8 million parliamentary speech records covering more than 20 languages, with a total size of 1.2 billion words. A teacher-student framework was adopted during its creation: large language models (e.g., GPT-4o) were used to automatically annotate policy topic labels, and multilingual encoder models were fine-tuned to improve annotation efficiency. ParlaCAP not only provides policy topic classification, but also includes rich speaker and political party metadata, and integrates the prediction results of the ParlaSent multilingual sentiment analysis model. As a comprehensive resource for comparative political science research, it can be used to analyze issues such as the distribution of political attention, emotional patterns, and gender differences in policy attention.



