Agent skill
claim-extraction
Extract structured claims, predictions, hints, and opinions from AI research content. Use when processing tweets, blog posts, substacks, or other content from AI researchers to identify substantive assertions about AI capabilities, limitations, and progress.
Install this agent skill to your Project
npx add-skill https://github.com/rickoslyder/HypeDelta/tree/main/.claude/skills/claim-extraction
SKILL.md
Claim Extraction Skill
Extract all substantive claims from AI research content. A claim is any assertion that:
- States something as true about AI capabilities, limitations, or progress
- Predicts future developments
- Hints at unreleased work
- Expresses a positioned opinion on the field's direction
- Critiques others' claims or work
Extraction Schema
For each claim, extract:
1. claimText
The claim in clear, standalone form. Paraphrase if needed for clarity.
2. claimType
fact: Assertion about current state ("GPT-4 can do X")prediction: Forward-looking ("By 2026, we'll have...")hint: Implies unreleased work ("We've been seeing interesting results with...")opinion: Positioned take ("I think scaling is/isn't sufficient")critique: Challenges others ("Marcus is wrong because...")question: Genuine uncertainty expressed ("I'm not sure if...")
3. topic
Primary topic category:
scaling: Scaling laws, compute, training efficiencyreasoning: LLM reasoning, chain-of-thought, planningagents: AI agents, tool use, autonomysafety: AI safety, alignment, controlinterpretability: Mechanistic interpretabilitymultimodal: Vision, audio, video modelsrlhf: RLHF, preference learning, Constitutional AIbenchmarks: Evals, benchmarks, capability measurementinfrastructure: Training infra, chips, hardwarepolicy: AI policy, regulation, governancegeneral: General AI commentary
4. stance
bullish: Optimistic about AI progress/capabilitiesbearish: Skeptical/pessimistic about AI progressneutral: Balanced or factual without clear stance
5. bullishness
Float from 0.0 (maximally bearish) to 1.0 (maximally bullish)
6. confidence
How confident does the author seem? (0.0-1.0)
- Hedging language: "might", "could", "I think", "possibly" → lower
- Certainty language: "will", "definitely", "it's clear that" → higher
7. timeframe (for predictions)
near-term: < 1 yearmedium-term: 1-3 yearslong-term: 3-10 yearsunspecified: No clear timeframenull: Not a prediction
8. evidenceProvided
strong: Cites data, papers, or detailed reasoningmoderate: Some reasoning but not rigorousweak: Assertion without supportappeal-to-authority: "Trust me, I work on this"
9. quoteworthiness
Is this claim notable enough to quote in a digest? (0.0-1.0)
Output Format
Return JSON:
{
"claims": [
{
"claimText": "The claim in clear form",
"claimType": "prediction",
"topic": "reasoning",
"stance": "bullish",
"bullishness": 0.8,
"confidence": 0.7,
"timeframe": "medium-term",
"evidenceProvided": "moderate",
"quoteworthiness": 0.6,
"relatedTo": ["o1", "chain-of-thought"],
"originalQuote": "Brief relevant quote if notable"
}
]
}
Guidelines
- Extract MULTIPLE claims from a single piece of content if present
- Don't over-extract - only substantive, meaningful claims
- A tweet saying "Interesting paper" is NOT a claim
- Look for IMPLICIT claims ("We've made a lot of progress" implies capability gains)
- Pay attention to WHO is speaking - lab researchers hinting at their own work is high signal
- Critics often make claims by contradiction ("X is wrong, therefore Y")
Author Context Matters
Consider the author's affiliation when assessing:
- Lab researchers (Anthropic, OpenAI, DeepMind): May hint at unreleased work
- Critics (Marcus, Chollet, Mitchell): Often make claims through critique
- Independent (Simon Willison, Jim Fan): Provide practitioner perspectives
Recommended Agent Skills
Expand your agent's capabilities with these related and highly-rated skills.
digest-generation
Generate a weekly AI intelligence digest from synthesized topic analyses and hype assessments. Use after synthesis and hype assessment to produce a readable, opinionated summary for sophisticated technical readers.
prediction-tracking
Track and evaluate AI predictions over time to assess accuracy. Use when reviewing past predictions to determine if they came true, failed, or remain uncertain.
topic-synthesis
Synthesize claims across multiple sources to identify consensus, disagreements, and emerging narratives on AI research topics. Use when you have claims from both lab researchers and critics on the same topic and need to understand where they agree, disagree, and what the overall hype level is.
content-filter
Filter and classify AI research content for relevance, topic, and author category. Use for bulk triage of raw content before detailed claim extraction.
hint-detection
Detect hints about unreleased AI research or capabilities from lab researcher communications. Use when analyzing tweets, posts, or interviews from people at major AI labs to identify signals about upcoming work.
hype-assessment
Assess overall hype levels across AI topics by comparing lab researcher enthusiasm against critic skepticism. Use after topic synthesis to identify which topics are overhyped, underhyped, or accurately assessed by the field.
Didn't find tool you were looking for?