---
title: "watch-video"
description: "Inspect videos with transcripts, timestamps, scene-aware frames, and optional transcription fallbacks."
canonical: https://agentpluginsdirectory.com/plugins/watch-video
last-updated: 2026-09-25
---

# watch-video
Inspect videos with transcripts, timestamps, scene-aware frames, and optional transcription fallbacks.
- Slug: watch-video
- Publisher: Nagarjuna Boddu
- Repository: https://github.com/heyNag/charms
- Manifest: packages/watch-video/plugin.json
- Version: 2026.8.10
- License: MIT
- Category (editorial): research
- Skills: 1 (watch-video)
- MCP servers: 0
- Stars: 1
- Repository created: 2026-06-19
- Repository last pushed: 2026-08-10
- Publisher type: User
- Listing: https://agentpluginsdirectory.com/plugins/watch-video
- Schema: https://agent-plugins.org/schemas/1.0.0/plugin.schema.json

## What watch-video does, in the publisher's words

watch-video is a local video inspection package for agents. It turns a URL or local video into a small evidence bundle that can include:

- metadata
- a focused audio clip when usable media is available
- transcript JSON and Markdown
- scene-aware frames with near-duplicate removal
- a concise report

Captions and metadata are probed before any media download. For URLs with usable captions, transcript detail skips media unless --timestamps pins cue frames. Frame selection is content-aware: scene changes by default, keyframe-first for fast skims with uniform fallback when fewer than four keyframes are available, and uniform sampling as the static-footage fallback. Sampled frame candidates have exact timestamps and pass through deduplication so held slides do not burn the frame budget; selected transcript-cue frames bypass deduplication.

From the project README, punctuation lightly normalized. Full text: https://raw.githubusercontent.com/heyNag/charms/HEAD/packages/watch-video/README.md

## Skills

- watch-video: Use when the user asks to inspect a YouTube URL, local video, screen recording, tutorial, demo, UI bug video, or visible/spoken video evidence.

Descriptions come from the frontmatter of each SKILL.md, punctuation lightly normalized.
