WildClawBench
收藏资源简介:
WildClawBench 是一个用于评估 AI 代理在真实环境中端到端任务执行能力的基准测试数据集。该数据集包含 60 个原创任务,覆盖六个主要类别:生产力流程、代码智能、社交互动、搜索与检索、创意合成和安全对齐。每个任务都设计用于测试 AI 代理在真实工作场景中的实际能力,如信息合成、多源聚合、代码库理解、多轮通信、网络搜索与本地数据协调、视频/音频处理等。数据集提供了一个隔离的 Docker 环境,包含 OpenClaw 实例和所有必要工具,确保任务的可重复性和隔离性。WildClawBench 旨在提供一个硬核、实用的评估平台,当前所有前沿模型的得分均低于 0.6,使得评分具有实际意义。
WildClawBench is a benchmark dataset for evaluating the end-to-end task execution capabilities of AI Agents in real-world environments. This dataset contains 60 original tasks covering six main categories: productivity workflows, code intelligence, social interaction, search and retrieval, creative synthesis, and safety alignment. Each task is designed to test the practical abilities of AI Agents in real workplace scenarios, such as information synthesis, multi-source aggregation, codebase comprehension, multi-turn communication, web search and local data coordination, video/audio processing, etc. The dataset provides an isolated Docker environment that includes an OpenClaw instance and all necessary tools, ensuring the reproducibility and isolation of the tasks. WildClawBench aims to provide a rigorous and practical evaluation platform, and the scores of all current state-of-the-art models are below 0.6, making the scoring results practically meaningful.




