Agent skill
video-caption-creation
Create on-screen text hooks and captions for short-form video clips. Built around the Complementarity Principle - on-screen text should ADD to what the viewer hears, not repeat it. Includes podcast clip workflow, hook categories, and Triple Word Score optimization.
Install this agent skill to your Project
npx add-skill https://github.com/majiayu000/claude-skill-registry/tree/main/skills/other/other/video-caption-creation
SKILL.md
Video Caption & On-Screen Hook Writer (v2)
On-screen text is the #1 visual element in short-form video. It does the PRIMARY work of stopping a scroll. A viewer sees the text and decides in <1 second whether to stop. This skill creates the text that makes them stop.
Previous version: SKILL_v1_archive.md (basic Triple Word Score, retired Feb 2026)
The Complementarity Principle
The central insight: On-screen text should NOT label or repeat the audio. It should ADD context that makes the audio land harder. The gap between what you read and what you hear creates curiosity.
This works exactly like title + thumbnail on YouTube: together they create a fuller picture than either alone. Andrew Muto: "You read the title, and then the thumbnail is offering you a little extra."
Bad (labeling): On-screen says "Kids know." → Audio is about teacher dissatisfaction. Vague label, no gap.
Good (complementary): On-screen says "A kindergartner can tell" → Audio reveals that 55% of teachers want to quit, and every morning a five-year-old walks into a classroom and senses it. The childlike framing + the adult data = productive tension.
The 6 Complementarity Patterns
| Pattern | On-Screen Text Does | Audio Does | Why It Works |
|---|---|---|---|
| Question → Answer | Asks a question | Delivers the answer | Open loop the viewer must close |
| Problem → Solution | Names a pain point | Provides the fix | Viewer self-identifies, stays for relief |
| Framework → Content | Names a strategy/concept | Explains what it means | "What IS that?" curiosity |
| Credential → Insight | Establishes authority | Delivers the revelation | Trust signal makes content land harder |
| Teaser → Payoff | Hints at what's coming | Completes the thought | Narrative tension drives completion |
| Reframe → Evidence | Challenges an assumption | Provides the proof | Cognitive dissonance demands resolution |
Worked Examples (Amar Kumar, KaiPod Learning)
Clip: "The Trapped Families Argument" Audio: "If we think that our public system's path to victory will be to trap families, you've lost the argument."
| Option | Pattern | Why It Works |
|---|---|---|
| "Schools trap families?" | Question → Answer | Question mark creates open loop. Broadest appeal. |
| "The school choice argument nobody makes" | Teaser → Payoff | Promises novelty. Viewer stays to hear the fresh angle. |
| "You're defunding yourself" | Reframe → Evidence | Provocative reframe. Viewer needs the logic. |
Clip: "Nobody Is Ever Lost" Audio: "In a big school, kids just feel lost. So they retreat to the screen. Then we ban the screen, and they're like, I'm still lost."
| Option | Pattern | Why It Works |
|---|---|---|
| "Why phone bans don't work" | Reframe → Evidence | Taps the biggest cultural conversation. Viewer thinks "wait, what?" |
| "Lost kids find screens" | Reframe → Evidence | Three words that reframe the phone debate. Unexpected word order. |
| "Ban the phone. Kid's still lost." | Problem → Solution | Mirrors the argument's rhythm. Punchy. |
| "80 schools. Zero phone problems." | Credential → Insight | Authority + counter-intuitive. Implies they solved it. |
Clip: "The Hidden Cost of a Dissatisfied Teacher" Audio: "55% of teachers want to leave... There's a cost every morning when a kindergartner walks into the classroom and says, I know this adult does not want to be here."
| Option | Pattern | Why It Works |
|---|---|---|
| "A kindergartner can tell" | Teaser → Payoff | Strongest gap. Childlike framing + adult data = productive tension. |
| "55% of teachers want to quit" | Reframe → Evidence | The stat does the work. "Wait, really?" |
| "Your kid's teacher hates their job" | Problem → Solution | Polarizing. Parents will react. Direct. |
| "Teachers aren't leaving for the money" | Reframe → Evidence | Challenges the dominant narrative. |
The First-3-Words Test
The first 3 words someone reads do 80% of the work. Before finalizing any hook, check:
What are the first 3 words?
- "A kindergartner can" → Curiosity (can what?)
- "Schools trap families" → Shock (they do?)
- "55% of teachers" → Specificity (tells me this is data)
- "Why phone bans" → Promise (I'm about to learn something)
Weak first 3 words:
- "Here's what happens" → Generic, could be anything
- "Check this out" → Zero information
- "In this clip" → Meta, breaks immersion
- "You need to" → Commanding without earning attention
Front-load the punch. Move the most surprising, specific, or provocative word as close to position 1 as possible.
Hook Categories
1. Polarizing Statement
Takes a side. Creates debate. Drives comments (= algorithm boost).
- "Schools trap families?"
- "Your kid's teacher hates their job"
- "Homework is a scam"
- Format:
[Institution/practice] [negative claim]
2. Counter-Intuitive Reveal
Surprises. Flips expectations. Challenges what the viewer assumed.
- "Why phone bans don't work"
- "Teachers aren't leaving for the money"
- "My son's first job sucked. Perfect."
- Format:
[Expected thing] [unexpected twist]
3. Direct Challenge
Confronts the viewer. Commands attention through provocation.
- "Your ADHD kid isn't broken"
- "Stop raising kids who wait for permission"
- "If your kid hates school, please watch this"
- Format:
[Imperative verb] [viewer's situation]
4. Curiosity Gap
Opens a loop the viewer must close by watching. Incomplete thought.
- "A kindergartner can tell"
- "The cost nobody talks about"
- "55% want out. Here's where they're going."
- Format:
[Partial information] [implied continuation]
5. Credential Lead
Establishes authority first, then promises the insight.
- "80 schools. Zero phone problems."
- "Harvard psychologist reveals..."
- "CEO of 80 microschools explains why..."
- Format:
[Authority signal] [promise] - When to use: Guest has impressive credentials. Research shows leading with expertise significantly boosts engagement (Diary of a CEO strategy).
6. Named Framework
Names a strategy or concept the viewer hasn't heard of.
- "The REPAIR strategy"
- "The Trojan Horse classroom"
- "The trapped families argument"
- Format:
The [name] [strategy/method/argument] - When to use: The speaker introduces a named concept, framework, or memorable phrase.
Podcast Clip On-Screen Hook Workflow
This is the step-by-step process for generating hooks from podcast transcript clips. Used by the podcast-production skill at Step 4.
Step 1: Identify the Hookable Moment
Scan the clip transcript for one of these triggers (SURP framework):
- Surprising fact or statistic ("55% of teachers want to leave")
- Unexpected quote or bold statement ("you've lost the argument")
- Relatable problem ("kids feel lost, so they retreat to the screen")
- Provocative opinion ("our path to victory will be to trap families")
The hookable moment is usually NOT the first sentence of the clip. Find the single line that would make someone stop scrolling if they saw it as a headline.
Step 2: Determine the Complementarity Pattern
Ask: What would ADD the most to what the viewer hears?
| If the audio... | Best pattern | On-screen text should... |
|---|---|---|
| Makes a bold claim | Reframe → Evidence | Challenge the assumption the claim disproves |
| Tells a story | Teaser → Payoff | Hint at the ending without revealing it |
| Cites data/stats | Question → Answer | Ask the question the stat answers |
| Introduces a framework | Framework → Content | Name the framework |
| Gives advice | Problem → Solution | Name the problem the advice solves |
| Shares credentials | Credential → Insight | Lead with the credential |
Step 3: Generate 3-5 Hook Options
For each clip, write 3-5 options across different categories. Variety matters - give the user a real choice.
Requirements per option:
- Hook text (3-8 words)
- Category label (from the 6 above)
- 1-line complementarity rationale (why this text + this audio = curiosity)
Step 4: Apply Quality Gates
For each option, check:
- First-3-Words Test: First 3 words carry the punch?
- McDonald's Test: Someone at McDonald's understands instantly?
- Scroll Test: Would I personally stop scrolling?
- Gap Test: Text + audio create a gap, not a repeat?
- Grandmother Test: My grandmother gets what this is about?
Step 5: Mark Recommended Pick
Select one recommended option and explain WHY in one sentence. The rationale should reference the complementarity - how the text and audio work together.
Identifying Hookable Moments in Transcripts
When scanning a full transcript (before clips are selected), look for:
Tier 1: Almost always hookable
- Specific statistics ("55% of teachers want to leave")
- Named concepts or frameworks the speaker coined
- Moments that contradict popular belief
- Emotional peaks (speaker's voice changes, pace quickens)
- Universal pain points ("every parent knows...")
Tier 2: Often hookable
- Personal confessions or vulnerable moments
- "I should probably not say this, but..."
- Origin stories (how they started, what went wrong)
- Direct statements that take a side
Tier 3: Rarely hookable (avoid)
- Agreement moments ("yeah, totally, I think so too")
- Background context / scene-setting
- Nuanced, heavily-caveated statements
- Inside baseball (jargon-heavy industry talk)
The whole transcript matters. The strongest moment might be at minute 48. Don't mine just the first half.
The Triple Word Score System
Four signals must align so algorithms AND humans immediately know "this is for me":
1. Audio Transcript (MOST IMPORTANT)
- What the speaker says out loud - algorithms auto-transcribe this
- Topic words must appear in first 10 seconds
- Core terminology repeated naturally throughout
2. On-Screen Text Hook
- Visual overlay using complementarity principle (above)
- NOT a repetition of the audio - an addition to it
- Lead with topic words in the hook
3. Caption Copy
- Post description with topic-relevant keywords
- Opens with a topic-relevant phrase
- Provides context the algorithm needs
4. Strategic Hashtags
- 10-12 total (optimal range)
- Broad → Mid → Specific → Niche → Audience → Platform
- Example: #Education #Parenting #Homeschool #Microschools #SchoolChoice #HomeschoolMom #Shorts
When all four align, the algorithm recognizes the topic immediately and serves it to the right people.
Caption Writing
Short-Form Platforms (Same Caption Everywhere)
Applies to: YouTube Shorts, Instagram Reels, TikTok, Facebook Reels
One caption per clip. Don't write platform-specific versions (waste of time, same audience).
Format:
[Verbatim quote or key insight from clip - 1-2 sentences]
[Context: who the speaker is and why it matters - 1 sentence]
#Hashtags (10-12)
Optional X Variant: Shorter, more conversational. Include @handles. Drop the context sentence.
Caption Quality Check
- Opens with the strongest line from the clip
- Includes guest name and title
- Hashtags span broad to niche
- Under 150 characters for the hook portion
- No emojis in body (rare exceptions for social captions)
Output Format
When this skill is invoked (standalone or from podcast-production), produce this per clip:
### [Clip Name]
**Timestamp range:** [MM:SS-MM:SS]
**On-screen hook options:**
1. "[Hook text]" *
2. "[Hook text]"
3. "[Hook text]"
4. "[Hook text]"
(* = recommended pick. List 3-4 options. Star goes on the strongest.)
**Caption (FB, TikTok, IG, LinkedIn):**
[Caption text]
**X variant:**
[Shorter caption with @handles]
Keep it clean. No category labels, no rationale paragraphs. The hooks should speak for themselves. Use the complementarity principle and quality gates internally when generating, but the output is just the options with a star.
Sub-Agent Prompt Template
When podcast-production invokes this skill at Step 4, use this prompt:
You are generating on-screen text hooks for podcast clips.
Episode: [Guest Name]
Working directory: [path to prep/]
Read the following files:
- SOURCE.md (full transcript for context)
- EDITOR_HANDOFF.md (clips already selected in Sections 3 & 4)
For EACH clip in Sections 3 and 4, generate 3-4 on-screen hook options.
Mark the recommended pick with an asterisk (*).
Output format per clip:
1. "Hook text" *
2. "Hook text"
3. "Hook text"
4. "Hook text"
No category labels. No rationale paragraphs. Just the options with a star.
Internal rules (apply these but don't show them in output):
- On-screen text ADDS to audio, never repeats it (Complementarity Principle)
- Apply the First-3-Words Test to every option
- Apply McDonald's Test (instantly understandable)
- Draw from the 6 hook categories for variety: Polarizing, Counter-Intuitive, Direct Challenge, Curiosity Gap, Credential Lead, Named Framework
- Hook text should be 3-8 words maximum
Write the hooks directly into the EDITOR_HANDOFF.md clip sections.
Exemplar Channels (Study These)
Podcast Clip Masters
- Diary of a CEO (@doac.clips) - Animated text, credential leads, emotional transitions. Opens with guest in chair + text hook. Text builds anticipation without giving away payoff.
- My First Million (clips channel) - Entrepreneurship hooks, "how to" + numbers format
- Huberman Lab (Essentials) - Dense science distilled to protocols. Direct promise hooks.
- Lex Fridman Clips - Topic-based text overlays for philosophical conversations
Education/Parenting Creators
- Dr. Becky Kennedy (@drbeckyatgoodinside, 3M) - Named strategy hooks ("The REPAIR Strategy"), problem → solution format
- Big Little Feelings (@biglittlefeelings, 3.6M) - "When your toddler refuses to..." pain point hooks
- Busy Toddler (@busytoddler, 2.4M) - Action-oriented "This LEGO hack will..."
- @thatcalteacherlife (508K) - Working parents + homeschool
- @deal_family (352K) - Second-gen homeschool mom
- @littlefenders (132K) - "Redefining learning + parenting / Educating kids with AI / Unschooling / ADHD Mom"
Key Patterns from Top Performers
- Text = Question, Audio = Answer (Dr. Becky)
- Text = Named Framework, Audio = Explanation (Dr. Becky)
- Text = Credential, Audio = Insight (DOAC)
- Text = Problem, Audio = Solution (Big Little Feelings)
- Text = Teaser, Audio = Payoff (DOAC)
Common Mistakes
- Labeling instead of complementing: "Kids know." is a label. "A kindergartner can tell" is a hook.
- Giving away the payoff in text: If the text reveals the full insight, there's no reason to watch.
- Generic first 3 words: "Here's what happens" could be anything. "55% of teachers" is specific.
- Only one hook per clip: Always generate 3-5. The first idea is rarely the best.
- Same category for every hook: If all 5 options are Curiosity Gap, you haven't explored enough.
- Hooks that need the clip to make sense: The hook must work for a silent, scrolling viewer who hasn't heard a word yet.
- Fancy vocabulary: "Pedagogical paradigm shift" fails the McDonald's test. "Schools are broken" passes.
Related Skills
podcast-production- Invokes this skill at Step 4 (on-screen hook generation)short-form-video- Full production workflow (this skill handles text only)text-content- For text-only social posts (LinkedIn, X, not video)youtube-title-creator- Title + thumbnail (same complementarity principle, different application)cold-open-creator- Cold opens use [SWOOSH] transitions, not hooks
Research References
Short-Form Video Text Hook Strategies.md- Comprehensive report on hook psychology, typography, retention- Andrew Muto notes:
Studio/_archive/Archive/Andrew Muto/- Complementarity, McDonald's test, 15-min rule, hook categories - Podcast-production skill improvement notes:
Studio/Podcast Studio/Amar-Kumar/prep/SKILL_IMPROVEMENT_NOTES.md
Rewritten Feb 6, 2026 after Amar Kumar podcast session. Incorporates complementarity principle, first-3-words test, 6 hook categories, SURP framework, and worked examples from real clips.
Recommended Agent Skills
Expand your agent's capabilities with these related and highly-rated skills.
agent-ops-spec
Manage specification documents in .agent/specs/. Use when user provides requirements, acceptance criteria, or feature descriptions that need to be tracked and validated against implementation.
agent-ops-state
Maintain .agent state files. Use at session start, after meaningful steps, and before concluding: read/update constitution/memory/focus/issues/baseline consistently.
agent-ops-spec
Manage specification documents in .agent/specs/. Use when user provides requirements, acceptance criteria, or feature descriptions that need to be tracked and validated against implementation.
agent-ops-testing
Test strategy, execution, and coverage analysis. Use when designing tests, running test suites, or analyzing test results beyond baseline checks.
agent-ops-testing
Test strategy, execution, and coverage analysis. Use when designing tests, running test suites, or analyzing test results beyond baseline checks.
agent-ops-state
Maintain .agent state files. Use at session start, after meaningful steps, and before concluding: read/update constitution/memory/focus/issues/baseline consistently.
Didn't find tool you were looking for?