PyMUSAS multilingual semantic annotation dataset
收藏资源简介:
该数据集由兰卡斯特大学团队构建,是首个针对USAS语义标注框架的多语言开放资源,包含银标准英文训练数据(约665万Tokens)及手动标注的中文评估数据集。数据源自高质量维基百科文档及特定领域文本(如芬兰咖啡网站、军事新闻),通过规则标注与人工校验结合生成。其核心价值在于解决多语言语义消歧任务中缺乏标注数据的问题,支持英语、芬兰语、威尔士语、爱尔兰语和中文的语义分析模型训练与评估。
This dataset, constructed by the team at Lancaster University, is the first multilingual open resource targeting the USAS semantic annotation framework. It includes silver-standard English training data (approximately 6.65 million Tokens) and a manually annotated Chinese evaluation dataset. The data is sourced from high-quality Wikipedia documents and domain-specific texts (e.g., Finnish coffee websites, military news), and is generated via a combination of rule-based annotation and manual verification. Its core value lies in addressing the shortage of annotated data for multilingual semantic disambiguation tasks, and supports the training and evaluation of semantic analysis models for English, Finnish, Welsh, Irish, and Chinese.




