---
title: "showdown"
description: "Minimalist Human-in-the-Loop & LLM Output Ranking Arena with preference convergence."
canonical: https://agentpluginsdirectory.com/plugins/showdown
last-updated: 2026-09-21
---

# showdown
Minimalist Human-in-the-Loop & LLM Output Ranking Arena with preference convergence.
- Slug: showdown
- Publisher: SgtPooki
- Repository: https://github.com/SgtPooki/showdown
- Manifest: plugin.json
- Version: 0.1.0
- License: MIT
- Category (editorial): other
- Skills: 1 (showdown)
- MCP servers: 0
- Stars: 0
- Repository created: 2026-09-18
- Repository last pushed: 2026-09-18
- Publisher type: User
- Listing: https://agentpluginsdirectory.com/plugins/showdown
- Schema: https://agent-plugins.org/schemas/1.0.0/plugin.schema.json

## What showdown does, in the publisher's words

> Minimalist Human-in-the-Loop & LLM Output Ranking Arena.

Showdown is a lightweight, keyboard-driven pairwise ranking and evaluation server for comparing outputs from LLMs, agents, or generative models.

It provides an active Elo matchmaking engine, rapid triage rating, real-time leaderboards, and direct export to standardized DPO (Direct Preference Optimization) training datasets.

- Multi-Modal Evaluations: Compares prose, markdown, code, JSON, and images.
- Active Elo Engine: Matchmaker actively pairs candidates with similar ratings and fewest evaluations for rapid convergence.
- Zero-Friction Keyboard UI: Hotkeys for instant voting (1 for A, 2 for B, T for Tie, S to Skip).
- Agent Integration: Simple Python SDK and REST API so any autonomous agent can launch a tournament, register candidates, and notify the user.
- DPO Dataset Export: Exports pairwise preferences directly to {prompt, chosen, rejected} JSONL for model alignment and fine-tuning.

From the project README, punctuation lightly normalized. Full text: https://raw.githubusercontent.com/SgtPooki/showdown/HEAD/README.md

## Skills

- showdown: Run interactive pairwise human-in-the-loop ranking and preference convergence on candidate outputs (text, code, markdown, images). Use to rank LLM responses, evaluate taglines or design concepts, capture user likes/dislikes, and iteratively evolve generation N+1 toward user preferences.

Descriptions come from the frontmatter of each SKILL.md, punctuation lightly normalized.
