inclusionAI/FinixDocBench
收藏资源简介:
FinixDocBench是一个金融领域文档解析基准数据集,专注于金融工作流程中常见但现有基准中代表性不足的文档解析条件。它包含742个页面样本,分为四个子集:FinixDigital(242个数字原生保险条款页面)、FinixPhoto(300个移动设备拍摄的医疗收据页面)、FinixHuge-Long(100个超长金融或保险页面)和FinixHuge-Table(100个大型密集表格页面)。数据集支持全页面Markdown解析、结构化布局解析(适用于FinixDigital和FinixPhoto)和超大页面可处理性评估。每个样本包括图像、Markdown文件和结构化JSON注释(部分子集)。数据集旨在评估OCR和文档解析系统在金融领域文档上的性能,特别是针对噪声文档、超大页面和复杂表格的鲁棒性。
FinixDocBench is a financial-domain document parsing benchmark that focuses on document parsing conditions common in real financial workflows but underrepresented in saturated clean-document benchmarks. It contains 742 page samples across four subsets: FinixDigital (242 digitally native insurance terms pages), FinixPhoto (300 mobile-captured medical receipt pages), FinixHuge-Long (100 ultra-long financial or insurance pages), and FinixHuge-Table (100 large dense table pages). The dataset supports full-page Markdown parsing, structured layout parsing (available for FinixDigital and FinixPhoto), and ultra-large page processability evaluation. Each sample includes an image, a Markdown file, and structured JSON annotations (for some subsets). It is intended for evaluating OCR and document parsing systems on financial documents, particularly for robustness on noisy documents, oversized pages, and complex tables.




