---
title: "bedrock-bench"
description: "Plan agent experiments, benchmark real tasks, compare saved results, diagnose failures, audit graders, and explore charts and tool replays."
canonical: https://agentpluginsdirectory.com/plugins/bedrock-bench
last-updated: 2026-10-06
---

# bedrock-bench
Plan agent experiments, benchmark real tasks, compare saved results, diagnose failures, audit graders, and explore charts and tool replays.
- Slug: bedrock-bench
- Publisher: OpenAI on AWS
- Repository: https://github.com/openai-on-aws/benchmarks-openai
- Manifest: plugins/bedrock-bench/plugin.json
- Version: 0.3.0
- License: MIT-0 AND CC-BY-SA-4.0
- Category (editorial): other
- Skills: 7 (audit-benchmark, benchmark-agent-tasks, compare-experiments, design-experiment, diagnose-failures, inspect-results, visualize-results)
- MCP servers: 0
- Stars: 3
- Repository created: 2026-06-26
- Repository last pushed: 2026-10-02
- Publisher type: Organization
- Listing: https://agentpluginsdirectory.com/plugins/bedrock-bench
- Schema: https://agent-plugins.org/schemas/1.0.0/plugin.schema.json

## What bedrock-bench does, in the publisher's words

Compare AI agents on real tasks and see how often they succeed, how long they take, and what each successful task costs. Run AWS CDK repairs, terminal tasks, and software-engineering benchmarks from a Codex conversation.

Install the plugin once, then select Bedrock Bench in chat and describe what you want to test.

From the project README, punctuation lightly normalized. Full text: https://raw.githubusercontent.com/openai-on-aws/benchmarks-openai/HEAD/plugins/bedrock-bench/README.md

## Skills

- audit-benchmark: Audit benchmark graders with correct solutions, alternate valid answers, malformed artifacts, and deliberate semantic mutants. Use to test false positives, false negatives, or AWS invariants before trusting benchmark results.
- benchmark-agent-tasks: Run AWS CDK repair, Terminal-Bench, SWE-bench, AWS-Bench, or starter agent tasks from chat; compare success, latency, token usage, and cost per successful task across Bedrock, OpenAI, OpenRouter, Codex, and OpenCode.
- compare-experiments: Compare compatible saved Bedrock Bench experiments using task-paired effects, task-cluster uncertainty, complete cost accounting, and chart-ready exports. Use for baseline versus candidate analysis of recorded results.
- design-experiment: Turn a benchmark question into a bounded, reproducible Bedrock Bench experiment plan with explicit conditions, runnable configs, task counts, and budget assumptions before execution.
- diagnose-failures: Inventory and triage failures across saved Bedrock Bench runs, with grounded status categories, missing attempts, cost coverage, and evidence citations. Use for recurring failure patterns or failure heatmaps; use inspect-results for one attempt and compare-experiments for comparable performance dif…
- inspect-results: Browse the Bedrock Bench run library, replay recorded tools, compare compatible runs, and explain outcomes using saved reports and per-attempt evidence.
- visualize-results: Create charts and diagrams from saved Bedrock Bench runs, including cost versus success, timing, token usage, compatible run history, and recorded tool sequences.

Descriptions come from the frontmatter of each SKILL.md, punctuation lightly normalized.
