Introduction MADCAT (Multilingual Automatic Document Classification Analysis and Translation) Phase 2 Training Set contains all training data created by the Linguistic Data Consortium
To use this dataset and respect for copyright, please cite the following paper: https://ieeexplore.ieee.org/abstract/document/9116896/ We present a new dataset that covers almost all the scenarios tha