CROHME+ 和 MathWriting+
收藏资源简介:
CROHME+ 和 MathWriting+ 数据集是为在线手写数学表达式识别(HMER)任务而创建的,它们提供了丰富的结构化注释,包括符号分割、分类和空间关系,这些数据集的创建旨在促进可解释的HMER研究。数据集包含374,000个数学表达式,为HMER任务提供了详尽的跟踪级别细节。这些数据集通过使用一个神经网络来自动地将LaTeX方程映射到原始跟踪,从而自动生成符号分割、分类和空间关系的注释。我们的结构识别系统生成一个完整的图结构,直接将手写跟踪链接到预测符号,从而实现透明的错误分析和可解释的输出。我们的结果挑战了结构方法过时的观念,证明了它们在高质量注释数据的支持下是可行的。
The CROHME+ and MathWriting+ datasets were developed for the task of Online Handwritten Mathematical Expression Recognition (HMER), providing rich structured annotations including symbol segmentation, classification and spatial relationships. These datasets were created to advance interpretable HMER research. The datasets contain 374,000 mathematical expressions, offering exhaustive trace-level details for the HMER task. These datasets automatically generate annotations for symbol segmentation, classification and spatial relationships by utilizing a neural network to map LaTeX equations to raw handwritten traces. Our structural recognition system generates a complete graph structure that directly links handwritten traces to predicted symbols, enabling transparent error analysis and interpretable outputs. Our results challenge the outdated notion that structural methods are obsolete, demonstrating their viability when supported by high-quality annotated data.




