CLiC-it Corpus
收藏资源简介:
CLiC-it语料库是一个收集了在意大利计算语言学会议(CLiC-it)的前十届会议中发表的693篇论文的元数据和文本内容的集合。该语料库旨在作为一个开放获取的结构化资源,用于研究趋势和意大利自然语言处理(NLP)社区的发展。它提供了关于作者、机构和主题的深入分析,并包含关于每篇研究论文的核心信息。语料库的设计允许进行纵向分析,并可以轻松扩展以包含未来会议的论文,从而支持对社区演变的持续监测。
The CLiC-it Corpus is a collection of metadata and full-text content for 693 papers published at the first ten editions of the Italian Conference on Computational Linguistics (CLiC-it). Developed as an open-access structured resource, it aims to support research on research trends and the developmental trajectory of the Italian Natural Language Processing (NLP) community. The corpus provides in-depth analyses of authors, affiliated institutions, and research topics, alongside core information for each individual research paper. Its design enables longitudinal analyses, and it can be easily extended to include papers from future conferences, thereby supporting continuous monitoring of the community’s evolution.

- 1Charting a Decade of Computational Linguistics in Italy: The CLiC-it CorpusCNR, Pisa - ItaliaNLP Lab · 2025年



