---
title: "hermes-jailbench"
description: "Jailbreak regression benchmark for LLM endpoints with repeatable known-pattern attacks and deterministic scoring"
canonical: https://agentpluginsdirectory.com/plugins/hermes-jailbench
last-updated: 2026-09-18
---

# hermes-jailbench
Jailbreak regression benchmark for LLM endpoints with repeatable known-pattern attacks and deterministic scoring
- Slug: hermes-jailbench
- Publisher: Hermes Labs
- Repository: https://github.com/hermes-labs-ai/hermes-jailbench
- Manifest: plugin.json
- Version: 0.2.1
- License: MIT
- Category (editorial): other
- Skills: 1 (hermes-jailbench)
- MCP servers: 0
- Stars: 3
- Repository created: 2026-04-17
- Repository last pushed: 2026-09-18
- Publisher type: Organization
- Listing: https://agentpluginsdirectory.com/plugins/hermes-jailbench
- Schema: https://agent-plugins.org/schemas/1.0.0/plugin.schema.json

## What hermes-jailbench does, in the publisher's words

hermes-jailbench is a jailbreak regression benchmark that runs a repeatable battery of known-pattern attacks against an Anthropic or OpenAI-compatible model endpoint and uses deterministic keyword heuristics to classify each response as refusal, partial, or compliance, so you can tell when a model or prompt update silently got less safe on attacks it used to refuse.

- "We changed the system prompt and now I need to know if refusals got weaker."
- "Our jailbreak testing lives in screenshots and anecdotes instead of something repeatable."
- "I want a no-key smoke test before I point real credentials at the model."
- "I need a known-pattern baseline before I claim a model is safer."

From the project README, punctuation lightly normalized. Full text: https://raw.githubusercontent.com/hermes-labs-ai/hermes-jailbench/HEAD/README.md

## Skills

- hermes-jailbench: Run the hermes-jailbench safety regression check against a model endpoint the user owns or is authorized to test, then summarize the report. Trigger when the user wants to confirm that a model, system prompt, or release change did not weaken refusals on a fixed set of known patterns, or wants a pas…

Descriptions come from the frontmatter of each SKILL.md, punctuation lightly normalized.
