OpenEarthAgent
收藏资源简介:
OpenEarthAgent 数据集是一个大规模、工具增强的地理空间推理语料库,旨在训练和评估多模态代理在结构化、多步地球观测(EO)任务上的能力。与传统专注于感知(分类、检测、分割)的遥感数据集不同,该数据集通过显式的工具交互支持可解释的多步推理,涵盖光学卫星图像、SAR图像、GIS矢量图层、地理参考栅格(GeoTIFF)以及光谱指数图层(如NDVI、NBR、NDBI等)。每个实例包括自然语言查询、多模态地理空间输入、结构化推理轨迹、显式工具调用及参数、中间工具观察结果和最终接地答案。该数据集适用于工具增强的LLM、地理空间推理、多模态代理、可解释的EO工作流程以及具有空间基础的结构化规划等研究方向。数据集统计显示,训练集包含14,538个实例和100,656个推理步骤,平均每个查询6.92步;测试集包含1,169个实例和7,064个推理步骤,平均每个查询6.04步,整个语料库的总推理步骤超过107K结构化思维-动作-观察转换。
The OpenEarthAgent dataset is a large-scale, tool-augmented geospatial reasoning corpus designed to train and evaluate multimodal agents on structured, multi-step Earth Observation (EO) tasks. Unlike traditional remote sensing datasets that focus on perception tasks such as classification, detection, and segmentation, this corpus supports interpretable multi-step reasoning via explicit tool interactions, covering optical satellite imagery, SAR imagery, GIS vector layers, georeferenced rasters (GeoTIFF), and spectral index layers (e.g., NDVI, NBR, NDBI, etc.). Each instance consists of a natural language query, multimodal geospatial inputs, structured reasoning trajectories, explicit tool calls and their corresponding parameters, intermediate tool observations, and final grounded answers. This dataset is applicable to research directions such as tool-augmented LLMs, geospatial reasoning, multimodal agents, interpretable EO workflows, and spatially grounded structured planning. Dataset statistics show that the training split contains 14,538 instances and 100,656 reasoning steps, averaging 6.92 steps per query; the test split contains 1,169 instances and 7,064 reasoning steps, averaging 6.04 steps per query, with the total reasoning steps across the entire corpus exceeding 107K structured thought-action-observation transitions.



