---
title: "agent-eval-tools"
description: "Experimental evaluation and improvement of agent skills and plugins, including improvement methods."
canonical: https://agentpluginsdirectory.com/plugins/agent-eval-tools
last-updated: 2026-10-05
---

# agent-eval-tools
Experimental evaluation and improvement of agent skills and plugins, including improvement methods.
- Slug: agent-eval-tools
- Publisher: peaceroad
- Repository: https://github.com/peaceroad/ai-dotfiles
- Manifest: plugins/agent-eval-tools/plugin.json
- Version: 0.1.0
- Category (editorial): other
- Skills: 1 (agent-improve)
- MCP servers: 0
- Stars: 0
- Repository created: 2026-07-17
- Repository last pushed: 2026-10-05
- Publisher type: User
- Listing: https://agentpluginsdirectory.com/plugins/agent-eval-tools
- Schema: https://agent-plugins.org/schemas/1.0.0/plugin.schema.json

## What agent-eval-tools does, in the publisher's words

An experimental Agent Plugins v1 package for evaluating and improving agent skills and plugins across their subject areas. Its single entry point, agent-improve, covers evaluation, comparisons, bounded improvement search, and applying that process to the improvement method itself.

- Diagnose a recurring skill failure and compare a repair with preservation cases.
- Test whether simplifying instructions preserves quality.
- Explore a capability goal even when no failure log exists.
- Remeasure a skill after a model change.
- Use retained evidence to compare a new improvement method with the current one.

Ordinary edits and design reviews can continue with existing design skills. This package adds experiment execution and evidence; it does not require a separate design plugin. Useful observations are retained selectively, and self-improvement is a separately scoped run rather than mandatory overhead after every task.

From the project README, punctuation lightly normalized. Full text: https://raw.githubusercontent.com/peaceroad/ai-dotfiles/HEAD/plugins/agent-eval-tools/README.md
