HermesBench

workflow reliability evals for personal AI agents

Last verified:

Visit HermesBench

What is HermesBench?

HermesBench is a runtime evaluation framework that benchmarks complete Hermes personal AI agent configurations, not just the underlying model. It evaluates the full agent setup including prompts, model/provider choices, tools, AgentSkills, memory systems, gateway behavior, delegation, safety mechanisms, latency, and stability. The current public baseline scores 78.2 across 27 personal-agent recipes with redacted traces available for inspection.

The tool provides 9 scored suites covering everyday personal-agent work including context management, calendar operations, web tasks, reports, communication, location services, travel planning, finance tasks, safety checks, and power-user integrations. Every published result links back to scenario definitions, public score axes, driver closure decisions, deterministic checks, and redacted trace timelines for full transparency.

HermesBench is designed for Hermes Agent users, developers building personal AI agents, and anyone who needs to evaluate workflow reliability across sessions, tool chains, and API failures. It uses a reliability-first scoring philosophy that penalizes lopsided scores, recognizing that a capable but unsafe agent or a safe but unhelpful agent is not actually good.

The platform offers public recipes where users can see prompts, redacted traces to inspect what happened during testing, and comprehensive methodology documentation explaining capability, reliability, and UX axes with documented limitations. Users can run benchmarks through a coding agent using the HermesBench skill or via Python API.

HermesBench pricing

Pricing model: Freemium

Free tier available with no credit card required. The public baseline is freely accessible with 27 workflow recipes and 9 scored suites. Full bundle runs are opt-in and cost more due to longer execution time. Python API default single-recipe path is available for free. Users can run one default scenario recipe for their current Hermes configuration at no cost. Profile and recipe submissions are free through the GitHub repository.

HermesBench pros

  • Evaluates complete agent configuration not just model performance
  • 27 user-like personal-agent workflow recipes included
  • 9 scored suites covering diverse personal-agent tasks
  • Redacted traces available for full transparency and inspection
  • 78.2 current public baseline with verifiable scores
  • Links to scenario definitions and score axes for every result
  • Deterministic checks ensure reproducible results
  • Covers everyday tasks: calendar, web, communication, travel, finance
  • Reliability-first scoring philosophy penalizes unsafe agents
  • Agent-driven quick start through coding agents like Codex or Claude
  • Python API available for automated benchmarking
  • Public recipes show exact prompts used for testing
  • Redacted profiles enable sharing successful configurations
  • Recipe submission system allows community contributions
  • Profile submission lets users share working setups
  • Simple public pathway: copy prompt to coding agent
  • No credit card required for basic usage

HermesBench cons

  • Currently in alpha stage needing early feedback
  • Full bundle runs take longer and cost more
  • Leaderboard not yet available as baseline is early
  • Limited to Hermes Agent configurations only
  • Requires coding agent setup for quick start pathway
  • Single baseline published so comparison data is limited
  • Opt-in for broader suites adds complexity
  • Redacted traces hide raw private payloads limiting full inspection
  • Documentation scattered across methodology document and site
  • Python API requires technical setup knowledge
  • No clear pricing for paid tiers or enterprise usage
  • Focus on personal agents may not suit enterprise workflows
  • Limited to 27 recipes may not cover all use cases
  • Scoring formulas detailed in separate methodology document
  • Driver closure decisions may be unclear to new users

Frequently asked questions about HermesBench

What is HermesBench?

HermesBench is a runtime evaluation framework that benchmarks complete Hermes personal AI agent configurations. It evaluates the full agent setup including prompts, model/provider, tools, AgentSkills, memory, gateway behavior, delegation, safety, latency, and stability—not just the underlying model performance.

How do I run a benchmark?

The public user pathway is intentionally simple: copy the provided prompt to Codex, Claude, or another coding agent. The agent loads the HermesBench skill and drives one scenario recipe first. Full bundle runs are opt-in. You can also use the Python API default single-recipe path, save artifacts, and summarize the score and main findings.

What is the current baseline score?

The current public baseline scores 78.2 across 27 personal-agent recipes with redacted traces you can inspect. This represents one early baseline, not a base-model leaderboard.

What workflows are covered?

The bundled catalog covers everyday personal-agent work including context management, calendar operations, web tasks, reports, communication, location services, travel planning, finance tasks, safety checks, and power-user integrations. There are 27 workflow recipes across 9 scored suites.

How is scoring calculated?

HermesBench uses six scoring axes: outcome reached, evidence/truthfulness, runtime/scope safety, responsiveness, task fulfillment, and communication quality. It is reliability-first but not capability-blind. Lopsided scores are penalized because a capable but unsafe agent or safe but unhelpful agent is not actually good. Detailed formulas are in the methodology document.

Can I inspect the benchmark results?

Yes. Every published result links back to scenario definitions, public score axes, driver closure decisions, deterministic checks, and redacted trace timelines. You can view redacted transcripts, tool timelines, checks, and judge reasoning without raw private payloads through the Traces tab.

How do I submit my profile or configuration?

Use the HermesBench skill to prepare your current Hermes profile/config as a public profile submission. Run one representative recipe first, package the redacted profile snapshot and score evidence, and follow the skill's guidance on what must be reviewed before opening a pull request.

Can I add new recipes to HermesBench?

Yes. Use the HermesBench skill to propose a new generic personal-agent recipe. Make the use case privacy-safe, driver/target agnostic, fixture-backed where possible, and include deterministic checks before preparing a pull request for community review.

What is the difference between recipes, profiles, and traces?

Recipes show what was tested—you can search by category, prompt, goal, and criteria. Profiles show what setup ran—review profile units, roles, and observed tools. Traces show what happened—open redacted transcripts, tool timelines, checks, and judge reasoning.

Is HermesBench suitable for enterprise use?

HermesBench focuses on personal AI agent configurations and everyday personal-agent work. It is designed for Hermes Agent users and developers building personal agents. The current alpha stage emphasizes workflow reliability evals for personal agents across sessions, tool chains, and API failures.

Categories

Use cases

Browse all AI tools on NeedAnAI