---
title: "mizan-eval"
description: "Run Mizan evaluations from an agent: score a single response or compare two responses against a metric template via the mizan CLI, reasoning over the -o json result. Use when you need to evaluate an asset/response agains"
canonical: https://agentpluginsdirectory.com/plugins/mizan-eval
last-updated: 2026-09-18
---

# mizan-eval
Run Mizan evaluations from an agent: score a single response or compare two responses against a metric template via the mizan CLI, reasoning over the -o json result. Use when you need to evaluate an asset/response against guidance and explain the verdict.
- Slug: mizan-eval
- Publisher: ghchinoy
- Repository: https://github.com/ghchinoy/mizan
- Manifest: plugins/mizan-eval/plugin.json
- Version: 0.1.0
- License: Apache-2.0
- Category (editorial): other
- Skills: 2 (run-eval-set, run-eval)
- MCP servers: 0
- Stars: 0
- Repository created: 2026-08-07
- Repository last pushed: 2026-09-17
- Publisher type: User
- Listing: https://agentpluginsdirectory.com/plugins/mizan-eval
- Schema: https://agent-plugins.org/schemas/1.0.0/plugin.schema.json

## Skills

- run-eval-set: Run a curated multi-concern Mizan eval-set (an EvalSet manifest) against a shared asset/response, interpret the weighted scorecard, per-member verdicts, the aggregate, the overall PASS/FAIL, and the gate, and honor the gate exit code for CI, driving the `mizan` CLI over `-o json`. Use when a user w…
- run-eval: Run a single (pointwise) or pairwise Mizan evaluation of a response/asset against a metric template and explain the verdict, driving the `mizan` CLI over `-o json` and reasoning over the parsed result. Use when a user asks to evaluate, score, grade, or A/B-compare a text response or media asset aga…

Descriptions come from the frontmatter of each SKILL.md, punctuation lightly normalized.
