Orionfold/hermes-brain-bench-v0.1
收藏资源简介:
Hermes Brain Bench v0.1是一个小型、多样化、基于分级评分标准的基准测试,用于在本地Spark部署中选取OpenAI兼容工具使用助手的代理大脑(最初针对Hermes Agent开发,但该测试套件与工具无关——任何支持OpenAI工具格式的运行器均可使用)。它通过10个提示(其中8个为核心提示,2个为条件提示)和7个字节确定性的种子固定装置,评估代理在工具调用、多步推理、诚实性、JSON格式遵守等方面的表现。该基准旨在回答单一流吞吐量基准无法解决的问题:在相同评分标准下,哪个本地服务通道实际上能产生更正确的代理?它提供了参考分数和详细评分方法,适用于比较不同本地服务通道的代理性能。
A small, diverse, graded-rubric benchmark for picking the agent brain behind a local-only Spark deployment of an OpenAI-compatible tool-using assistant (developed against Hermes Agent, but the suite is harness-agnostic — any OpenAI-tool-format runner works). It measures agent performance through 10 prompts (8 core, 2 conditional) and 7 bytes-deterministic seeded fixtures, focusing on tool calling, multi-step reasoning, honesty, JSON format adherence, and more. The bench answers the question: which local serving lane actually produces the more correct agent under the same rubric, run-to-run? It includes reference scores and detailed scoring methodology for comparing agent performance across different local serving lanes.




