Inv3D: a high-resolution 3D invoice dataset for template-guided single-image document unwarping - Test split
收藏资源简介:
Numerous business workflows involve printed forms, such as invoices or receipts, which are often manually digitalized to persistently search or store the data. As hardware scanners are costly and inflexible, smartphones are increasingly used for digitalization. Here, processing algorithms need to deal with prevailing environmental factors, such as shadows or crumples. Current state-of-the-art approaches learn supervised image dewarping models based on pairs of raw images and rectification meshes. The available results show promising predictive accuracies for dewarping, but generated errors still lead to sub-optimal information retrieval. In this paper, we explore the potential of improving dewarping models using additional, structured information in the form of invoice templates. We provide two core contributions: (1) a novel dataset, referred to as Inv3D, comprising synthetic and real-world high-resolution invoice images with structural templates, rectification meshes, and a multiplicity of per-pixel supervision signals and (2) a novel image dewarping algorithm, which extends the state-of-the-art approach GeoTr to leverage structural templates using attention. Our extensive evaluation includes an implementation of DewarpNet and shows that exploiting structured templates can improve the performance for image dewarping. We report superior performance for the proposed algorithm on our new benchmark for all metrics, including an improved local distortion of 26.1 %. We made our new dataset and all code publicly available at https://felixhertlein.github.io/inv3d.
诸多商业流程中都会用到打印表单,例如发票或收据,这类表单通常需要人工进行数字化处理,以便持久化检索与存储其中的数据。由于硬件扫描仪成本高昂且灵活性不足,智能手机正愈发广泛地被用于表单数字化工作。在此场景下,处理算法需要应对各类常见的环境干扰因素,例如阴影与褶皱。当前主流的先进技术基于原始图像与校正网格(rectification meshes)的配对样本,训练有监督的图像去形变(image dewarping)模型。现有研究结果显示,该类模型的形变校正预测精度可观,但仍会产生校正误差,进而导致信息检索效果未达最优。本文旨在探索利用发票模板形式的结构化附加信息,以优化图像去形变模型性能的可行性。本文主要包含两项核心贡献:其一,构建了一款全新的数据集Inv3D,该数据集涵盖合成与真实场景下的高分辨率发票图像,附带结构化模板、校正网格以及多类逐像素监督信号;其二,提出了一种全新的图像去形变算法,该算法通过注意力机制利用结构化模板,对当前主流先进方法GeoTr进行了扩展。我们开展了全面的评估实验,其中包含对DewarpNet的复现,实验结果验证了利用结构化模板可有效提升图像去形变任务的性能。在我们构建的全新基准数据集上,所提算法在所有评测指标上均展现出更优性能,其中局部畸变校正指标提升了26.1%。我们已将全新数据集与全部代码公开至https://felixhertlein.github.io/inv3d。




