---
title: "orq"
description: "Agent skills for building, deploying, evaluating, and monitoring LLM pipelines on the orq.ai platform."
canonical: https://agentpluginsdirectory.com/plugins/orq
last-updated: 2026-09-28
---

# orq
Agent skills for building, deploying, evaluating, and monitoring LLM pipelines on the orq.ai platform.
- Slug: orq
- Publisher: orq.ai
- Repository: https://github.com/orq-ai/assistant-plugins
- Manifest: plugin.json
- Version: 3.4.0
- License: MIT
- Category (editorial): agent-tooling
- Skills: 17 (create-skill, evaluatorq, orq-analyze-traces, orq-build-agent, orq-build-evaluator, orq-cli, orq-compare-agents, orq-evaluator-alignment, orq-generate-synthetic-dataset, orq-improve-agent, orq-invoke-deployment, orq-manage-skills, orq-red-team, orq-run-experiment, orq-setup-observability, orq-shared, orq-simulate-agent)
- MCP servers: 0
- Stars: 6
- Repository created: 2026-03-04
- Repository last pushed: 2026-09-28
- Publisher type: Organization
- Listing: https://agentpluginsdirectory.com/plugins/orq
- Schema: https://agent-plugins.org/schemas/1.0.0/plugin.schema.json

## What orq does, in the publisher's words

Agent Skills for the full Build → Evaluate → Optimize lifecycle of LLM pipelines on orq.ai.

Skills are multi-step workflows that require reasoning (e.g. build an agent, run an experiment);

Commands are quick actions for immediate results (list traces, show analytics).

From the project README, punctuation lightly normalized. Full text: https://raw.githubusercontent.com/orq-ai/assistant-plugins/HEAD/README.md

## Skills

- create-skill: Build or update an agent skill from an API, CLI, or MCP surface. Use when the user wants to document a new capability surface as a skill, or when an existing skill's contract is stale and needs re-verification. Do NOT use when editing an existing skill without re-probing the surface, or when docume…
- evaluatorq: Write and run evaluatorq evaluation scripts (Python or TypeScript) for a single agent or deployment, custom scorers, built-in evaluators, and dataset-driven evaluation. For CLI workflows, use the companion skills: `orq-red-team` for `eq redteam` adversarial testing and `orq-simulate-agent` for `eq…
- orq-analyze-traces: Analyze a live agent, deployment, or local agent from its production traces, relay its configuration and terminal states, then build a failure taxonomy by open coding and axial coding, and write it to an error-analysis file other skills read. Use when debugging agent or pipeline quality, when you…
- orq-build-agent: Design, create, and configure orq.ai Agents with tools, instructions, knowledge bases, and memory stores. Use when building new agents, attaching KBs or memory, writing system instructions, selecting models, or setting up RAG pipelines. Do NOT use for debugging existing agents (use orq-analyze-trac…
- orq-build-evaluator: Create validated LLM-as-a-Judge evaluators following best practices, binary Pass/Fail judges by default, plus numeric and categorical, all validated against human labels for measuring specific failure modes. Use when you need to automate quality checks, build guardrails, or measure a specific fail…
- orq-cli: Drive the `orq` command-line interface: check the install, authenticate, select a workspace, and run read and write commands against any orq.ai resource (traces, agents, deployments, evals, prompts, datasets, projects, skills). Use when a task needs shell access to orq.ai, when a script or CI job…
- orq-compare-agents: Run cross-framework agent comparisons using evaluatorq from orqkit, compares any combination of agents (orq.ai, LangGraph, CrewAI, OpenAI Agents SDK, Vercel AI SDK) head-to-head on the same dataset with LLM-as-a-judge scoring. Use when comparing agents, benchmarking, or wanting side-by-side evalua…
- orq-evaluator-alignment: Align, calibrate, or improve an existing LLM-as-a-judge (orq evaluator) so its verdicts match human judgment, boolean, categorical, or numeric judges. Use when the user wants to "align my evaluator", "improve my eval", "my judge keeps changing its mind", "find ambiguous cases", or "annotate an eva…
- orq-generate-synthetic-dataset: Generate and curate evaluation datasets: structured generation via dimensions-tuples-NL, quick from description, expansion from existing data, plus dataset maintenance through deduplication, rebalancing, and gap-filling. Use when creating eval data, expanding test coverage, or cleaning datasets. D…
- orq-improve-agent: Improve an underperforming orq agent, deployment, or local agent, rewrite its instructions against a structured prompting framework, or move a configuration knob, grounded in the error-analysis file orq-analyze-traces writes. Use when a prompt needs improvement, when a config knob is wrong (trunca…
- orq-invoke-deployment: Invoke orq.ai deployments, agents, and models via the Python SDK or HTTP API. Use when a user wants to call a deployment with prompt variables, invoke an agent in a conversation, or call a model directly through the AI Router. Do NOT use for creating or editing deployments/agents (use orq-improve-a…
- orq-manage-skills: Manage orq.ai Skills (the platform entity, formerly called Snippets) end-to-end, list, get, create, update, and delete Skills, plus authoring guidance (display name, description, tags, project scoping, path placement), and how Skills get consumed (the `{{skill.<display_name>}}` template placeholde…
- orq-red-team
- orq-run-experiment
- orq-setup-observability
- orq-shared
- orq-simulate-agent

Descriptions come from the frontmatter of each SKILL.md, punctuation lightly normalized.
