---
title: "hopper"
description: "LLM, text-to-speech and speech-to-text for voice agents and AI agents: set Hopper up in a voice-agent project, measure time to first token on its own prompt, find what slows it down, and speak or transcribe audio."
canonical: https://agentpluginsdirectory.com/plugins/hopper
last-updated: 2026-10-01
---

# hopper
LLM, text-to-speech and speech-to-text for voice agents and AI agents: set Hopper up in a voice-agent project, measure time to first token on its own prompt, find what slows it down, and speak or transcribe audio.
- Slug: hopper
- Publisher: Hopper
- Repository: https://github.com/hopper-inc/plugins
- Manifest: hopper/plugin.json
- Version: 1.2.0
- License: MIT
- Category (editorial): other
- Skills: 5 (hopper-benchmark, hopper-diagnose, hopper-integrate, hopper-speak, hopper-transcribe)
- MCP servers: 1 (hopper)
- Stars: 0
- Repository created: 2026-09-30
- Repository last pushed: 2026-10-01
- Publisher type: Organization
- Listing: https://agentpluginsdirectory.com/plugins/hopper
- Schema: https://agent-plugins.org/schemas/1.0.0/plugin.schema.json

## What hopper does, in the publisher's words

LLM inference for voice agents, and a voice and ears for any agent. Hopper serves gemma-4-31b behind an OpenAI-compatible API built for low time to first token, plus text-to-speech and speech-to-text. An agent can start using it on its own: the first run registers a key for it with no human step, and the user can later move that key to their Hopper account. Works in Claude Code, Codex, ChatGPT, Claude and Cursor, and anywhere Agent Skills run.

In Claude Code they are also commands, such as /hopper:hopper-integrate.

The plugin connects the Hopper MCP server at https://withhopper.com/mcp, for hosts without a shell such as ChatGPT and Claude:

From the project README, punctuation lightly normalized. Full text: https://raw.githubusercontent.com/hopper-inc/plugins/HEAD/hopper/README.md

## Skills

- hopper-benchmark: Measure a voice agent''s LLM time to first token on Hopper with its own system prompt and tools over a simulated 10-turn call, and report first-turn and later-turn latency and the prompt-cache hit rate. Changes no code. Use when the user asks "how fast would my agent be on Hopper", "benchmark TTFT"…
- hopper-diagnose: Find what delays a voice agent''s first spoken word: timestamps or IDs at the top of the prompt, a new client per call, HTTP/1.1, no warm-up, thinking left on, tools named but not sent. Ranks fixes by the latency they recover. Works with any LLM provider, needs no key, and edits nothing until the u…
- hopper-integrate: Switch an existing voice agent''s LLM to Hopper, an OpenAI-compatible endpoint built for low time to first token. Gets a key for the project with no human step, benchmarks the agent''s own prompt and tools, and edits code only after the user says go. Use when the user says "make my agent respond fa…
- hopper-speak: Turn text into speech with Hopper text-to-speech: play it aloud, save a WAV file, or preview and pick a voice. Needs no setup: the first use registers a key for the agent itself. Use when the user says "read this aloud", "say this out loud", "generate a voiceover", "make an audio file of…", "TTS" o…
- hopper-transcribe: Transcribe speech in an audio or video file with Hopper speech-to-text, with word timestamps. Takes WAV, MP3, M4A, voice memos and video (decoded with ffmpeg). Needs no setup: the first use registers a key for the agent itself. Use when the user says "transcribe this", "what does this recording say…

Descriptions come from the frontmatter of each SKILL.md, punctuation lightly normalized.

## MCP servers

- hopper: transport: streamable-http; url: https://withhopper.com/mcp

Read from the plugin's own mcp.json. Environment variable names only, never values.
