docmath-eval-failures-200
收藏资源简介:
DocMath-Eval Failures 200 是一个精选的基准数据集,包含200个具有挑战性的金融数学问题,这些问题由领先的AI模型未能正确回答。该数据集旨在评估AI代理在金融文档上的数值推理能力。数据集包含多个配置,包括默认配置(包含问题和真实答案)、无答案配置(仅包含问题和上下文,用于公平评估)、失败配置(原始Gemini 2.5 Flash失败数据)、结果配置(所有代理预测和评分)、排行榜配置(每次运行的摘要统计)和分拆配置(每次运行的详细分拆)。数据集适用于问答和文本生成任务,特别适合用于评估AI代理在复杂金融数学问题上的表现。数据集还提供了详细的评估结果和排行榜,展示了不同代理和模型的表现。
DocMath-Eval Failures 200 is a curated benchmark dataset containing 200 challenging financial mathematics problems that were incorrectly answered by state-of-the-art AI models. This dataset is designed to evaluate the numerical reasoning capabilities of AI Agents when processing financial documents. The dataset includes multiple configurations: the default configuration (containing questions and ground-truth answers), the no-answer configuration (containing only questions and contexts for fair evaluation), the failures configuration (raw failure data from Gemini 2.5 Flash), the results configuration (all agent predictions and scoring results), the leaderboard configuration (summary statistics for each run), and the split configuration (detailed per-run splits). The dataset is suitable for question answering and text generation tasks, and is particularly well-suited for evaluating the performance of AI Agents on complex financial mathematics problems. Additionally, the dataset provides detailed evaluation results and leaderboards that showcase the performance of various agents and models.



