---
title: "save-toolkit"
description: "Application-engineering and site-reliability agents and reusable skills."
canonical: https://agentpluginsdirectory.com/plugins/save-toolkit
last-updated: 2026-09-25
---

# save-toolkit
Application-engineering and site-reliability agents and reusable skills.
- Slug: save-toolkit
- Publisher: latent-sre
- Repository: https://github.com/latent-sre/save-toolkit
- Manifest: plugin.json
- Version: 0.50.0
- License: MIT
- Category (editorial): other
- Skills: 29 (agent-authoring, akamai-edge, backend-craft, ci-actions, database-reliability, eng-ladder, frontend-craft, gcp-ops, grafana, incident-investigation, obs-alerting, obs-dashboards, obs-logs, obs-metrics, obs-pipeline, obs-traces, operational-learning, operator-cli, pcf-deploy, pcf-ops, postmortem, production-change-gate, python-craft, resilience-analysis, root-cause, runbook, service-lifecycle, stack-profile, toil-reduction)
- MCP servers: 0
- Stars: 0
- Repository created: 2026-06-15
- Repository last pushed: 2026-09-24
- Publisher type: User
- Listing: https://agentpluginsdirectory.com/plugins/save-toolkit
- Schema: https://agent-plugins.org/schemas/1.0.0/plugin.schema.json

## What save-toolkit does, in the publisher's words

A Claude Code plugin that helps a human SRE do their job on this team's stack: PCF (through Apps Manager), Splunk, Wavefront and PCF App Metrics, Grafana, Akamai, with Cloud Run migration guidance. The GCP landing runtime remains undecided in stack-profile. It fits together in three layers. You own the work: the incident, the change, the runbook. An advisor thinks with you: incident-investigation asks what to check next, says what each result means, and tells you when to mitigate. Agents are your helpers, dispatched by you or the invoking workflow for bounded jobs: sre-assistant performs a read-only lookup or investigation, observability-engineer tunes an alert, scribe writes the runbook afterward. The skills serve you and the agents alike: the same logs skill hands you a paste-ready Splunk search and hands the sre-assistant agent the method to build one, which is why a PCF check is always the Apps Manager view with the cf command beside it.

From the project README, punctuation lightly normalized. Full text: https://raw.githubusercontent.com/latent-sre/save-toolkit/HEAD/README.md

## Skills

- agent-authoring: Create, repair, or security-review LLM-facing prompts, agents, skills, tool descriptions, graders, bounded Loop Engineering for evaluation/verification, and agent roster/delegation graphs. Triggers: 'write me an agent/skill/prompt', 'my skill fires too often', 'the output is the wrong shape', 'is t…
- akamai-edge: Akamai edge work in three lanes: triage (edge vs origin, Reference # error strings, cache status, WAF denials, DataStream 2), delivery config (Property Manager versions, staging-first activation, fast fallback), and mPulse RUM (network-side vs app-side slowdowns). Triggers: 'is it the CDN or the or…
- backend-craft: Build or change an API or backend service, HTTP endpoints, workers, schedulers, the service behind a UI, and consume third-party APIs safely (clients, SDK wrappers, sync jobs, webhooks), including our platform/obs APIs. Triggers: 'add an endpoint', 'wrap X behind an API', 'write a client for Y'. No…
- ci-actions: Review, design, troubleshoot, and optimize GitHub Actions workflows: fast feedback, caching, test matrices, reusable jobs, reproducible artifacts, and secure delivery. Includes advice-only workflow reviews when the caller asks for findings without edits. Triggers: 'set up CI', 'speed up this pipeli…
- database-reliability: Diagnose and improve data-layer reliability: slow queries, lock contention, replication lag, connection pools, schema migrations, and recovery evidence. Triggers: 'this query is slow', 'plan this schema migration', 'the connection pool is exhausted'. Not for app-side triage (pcf-ops), burn alerts (…
- eng-ladder: Select the engineering altitude for implementation, design, review, or growth feedback when work may span components, teams, migrations, or hard-to-reverse choices. Triggers: 'how rigorous should this be', 'review this at the principal level', 'is this a design doc or just a PR'. A scoped change wi…
- frontend-craft: Build or change a web UI, pages, dashboards-as-app-features, forms, admin panels, from a single page to a full SPA, including serving it on PCF. Owns UI-layer TypeScript/React idiom: component state, interaction, accessibility, resilience UX. Triggers: 'build a UI for', 'add a page/form/table', 'ma…
- gcp-ops: Investigate application-side GCP failures during the migration, Cloud Run services and revisions, gcloud logging reads, what-changed correlation against revision deploys, and the project-vs-platform boundary. Triggers: 'the Cloud Run service is 503ing', 'read the GCP logs', 'container failed to lis…
- grafana: Operate Grafana: find, explain, create, and edit dashboards; inspect alert state and notification paths; create or update Grafana-managed alert rules, and manage temporary silences. Triggers: 'explain this Grafana dashboard', 'create a Grafana alert', 'silence this alert', 'edit this Grafana dashbo…
- incident-investigation: Helps a human SRE investigate a live incident, understand evidence, and choose the next useful step. Use for new pages, ongoing troubleshooting, interpreting supplied logs, graphs, metrics, traces or alerts, comparing mitigation options and recommending what to do, checking recovery, and preparing…
- obs-alerting: Design alerting that pages on symptoms: SLIs/SLOs and multi-window burn rates, Splunk saved-search alerts, Moogsoft correlation, and ThousandEyes synthetics. Triggers: 'define an SLO', 'this alert is too noisy', 'what should page', 'design a synthetic check'. Not for queries (obs-metrics, obs-logs)…
- obs-dashboards: Design dashboards around the on-call reader's questions: service health, golden signals, useful panels, units, comparisons, missing-data presentation, and drill-downs. Triggers: 'design a dashboard', 'what should we dashboard', 'which panels do we need', 'make this dashboard easier to read'. Grafan…
- obs-logs
- obs-metrics
- obs-pipeline
- obs-traces
- operational-learning
- operator-cli
- pcf-deploy
- pcf-ops
- postmortem
- production-change-gate
- python-craft
- resilience-analysis
- root-cause
- runbook
- service-lifecycle
- stack-profile
- toil-reduction

Descriptions come from the frontmatter of each SKILL.md, punctuation lightly normalized.
