Agent skill
kvpress
kvpress (NVIDIA) KV-cache compression for HuggingFace LLMs. Use when: kvpress imports, compression_ratio, press(model) context managers, StreamingLLMPress, SnapKVPress, ExpectedAttentionPress, TOVAPress, KnormPress, KV-cache eviction, token pruning during generation, or attention sink methods.
Install this agent skill to your Project
npx add-skill https://github.com/JoaquinCampo/Skills/tree/main/kvpress
SKILL.md
kvpress — KV-Cache Compression for LLMs
kvpress is an NVIDIA library that compresses the KV cache of HuggingFace transformers models during generation, reducing memory usage at the cost of potential quality degradation.
- Repository: https://github.com/NVIDIA/kvpress
- Paper: https://arxiv.org/abs/2510.00636v1
- Version: 0.5.1+ (requires transformers v5+)
- License: Apache 2.0
Core Concept
A "press" is a callable object that wraps a model as a context manager. Inside the context, forward hooks on every attention layer intercept the KV cache after prefilling and prune it according to the press's strategy. Generation then proceeds with the compressed cache.
from kvpress import StreamingLLMPress
press = StreamingLLMPress(compression_ratio=0.5)
with torch.no_grad(), press(model):
outputs = model.generate(**inputs, max_new_tokens=256)
compression_ratio — The #1 Gotcha
compression_ratio = fraction of KV pairs to REMOVE (not keep)
| compression_ratio | Effect |
|---|---|
0.0 |
No compression (keep 100%) |
0.5 |
Remove 50%, keep 50% |
0.875 |
Remove 87.5%, keep 12.5% |
1.0 |
Invalid (assertion fails) |
Internal calculation: n_kept = int(seq_len * (1 - compression_ratio))
If your code uses "fraction to keep" semantics, convert: kvpress_ratio = 1.0 - keep_fraction
Context Manager Mechanics
When you call press(model):
- Validates model architecture (warning if unsupported, not an error)
- Registers
forward_hook(with_kwargs=True)on everymodel.model.layers[i].self_attn - Yields control — your
model.generate()runs here - Removes all hooks on context exit
The hooks fire only during prefill (when q_len == k_len). During autoregressive generation, hooks are still registered but skip compression. This means:
- Compression is a one-time operation at the start of generation
- All generated tokens see the same compressed cache
- There is no ongoing compression during token-by-token generation (unless using DecodingPress)
Safe to use with:
output_scores=True— hooks operate on KV cache, not logitsreturn_dict_in_generate=True— no interferencedo_sample=False(greedy) ordo_sample=True(sampling)StoppingCriteria— works normally
Supported Models
SUPPORTED_MODELS = (
LlamaForCausalLM, # Llama 2, 3, 3.1, 3.2
MistralForCausalLM, # Mistral 7B, etc.
Phi3ForCausalLM, # Phi-3
Qwen2ForCausalLM, # Qwen2, Qwen2.5
Qwen3ForCausalLM, # Qwen3
Gemma3ForConditionalGeneration, # Gemma 3
)
The check is a warning, not a hard block. Models with model.model.layers[].self_attn structure may work even if not listed.
Class Hierarchy
BasePress (dataclass, context manager)
├── ScorerPress (score-based pruning, has compression_ratio)
│ ├── StreamingLLMPress — position-based: keep sinks + recent
│ ├── SnapKVPress — attention of recent tokens
│ ├── KnormPress — key vector L2 norms
│ ├── ExpectedAttentionPress — predicted future attention
│ ├── TOVAPress — last-token attention weight
│ ├── ObservedAttentionPress — full prefill attention (needs eager)
│ ├── RandomPress — random baseline
│ ├── KeyDiffPress — key distinctiveness
│ ├── LagKVPress — lag-relative information
│ ├── CURPress — leverage scores
│ ├── KVzapPress — learned surrogate (needs HF weights)
│ ├── QFilterPress — learned filters (needs HF weights)
│ ├── LeverageScorePress — statistical leverage via Cholesky
│ ├── NonCausalAttnPress — non-causal chunked attention
│ ├── CompactorPress — blends leverage + non-causal attn
│ ├── PyramidKVPress — extends SnapKV
│ └── CriticalKVPress — two-stage with value norms
├── ThinKPress — dimension compression (channels, not sequence)
├── SimLayerKVPress — layer-adaptive (lazy layer detection)
├── DuoAttentionPress — head-adaptive (retrieval vs streaming)
├── FinchPress — prompt-guided, delimiter-based
├── KVzipPress — context reconstruction (2-3x overhead)
├── FastKVzipPress — learned gates
└── Wrappers:
├── ComposedPress — chains multiple presses
├── AdaKVPress — head-wise adaptive (wraps ScorerPress)
├── ChunkPress — chunk-wise uniform compression
├── ChunkKVPress — semantic chunk selection
├── BlockPress — block-wise iterative
├── PerLayerCompressionPress — per-layer ratios
├── KeyRerotationPress — RoPE fix after pruning
├── DecodingPress — compression during decoding (experimental)
├── PrefillDecodingPress — separate prefill + decoding strategies
└── DMSPress — threshold-based adaptive
For detailed per-press documentation (parameters, papers, requirements), read:
`references/press-catalog.md`
## Quick Reference: Choosing a Press
### No special setup needed (just compression_ratio):
| Press | Strategy | Attention needed? | Best for |
|---|---|---|---|
| StreamingLLMPress | Keep sinks + recent tokens | No | Simple baseline, predictable behavior |
| SnapKVPress | Recent tokens' attention patterns | No (computes own) | Good general-purpose quality |
| KnormPress | Key vector norms | No | Fast, no attention compute |
| ExpectedAttentionPress | Predicted future attention | No | Best quality (NVIDIA's method) |
| TOVAPress | Last token's attention | Optional | Lightweight attention-based |
| RandomPress | Random eviction | No | Degradation baseline |
| KeyDiffPress | Key distinctiveness | No | Unique key preservation |
### Needs `attn_implementation="eager"`:
| Press | Why |
|---|---|
| ObservedAttentionPress | Uses full prefill attention matrix |
### Needs pre-trained weights from HF Hub:
| Press | Weights from |
|---|---|
| QFilterPress | `nthngdy/` (not all models) |
| KVzapPress | `nvidia/KVzap-{type}-{model}` |
| FastKVzipPress | Per-model on HF Hub |
| ExpectedAttentionStatsPress | Pre-computed query stats |
### Special context managers (incompatible with ComposedPress):
| Press | Why |
|---|---|
| KVzipPress | Multi-pass, 2-3x overhead |
| FastKVzipPress | Own `__call__` implementation |
| AdaKVPress | Uses attention_patch mechanism |
## Common Patterns
### Basic usage
```python
from kvpress import SnapKVPress
press = SnapKVPress(compression_ratio=0.5)
with torch.no_grad(), press(model):
out = model.generate(**inputs, max_new_tokens=512)
Baseline (no compression) — use nullcontext
from contextlib import nullcontext
ctx = press(model) if press is not None else nullcontext()
with torch.no_grad(), ctx:
out = model.generate(**inputs, max_new_tokens=512)
Composing presses
from kvpress import ComposedPress, SnapKVPress, ThinKPress
press = ComposedPress([
SnapKVPress(compression_ratio=0.3),
ThinKPress(key_channel_compression_ratio=0.2),
])
# Effective keep ratio = (1-0.3) * (1-0.2) = 0.56
Head-wise adaptive compression
from kvpress import AdaKVPress, SnapKVPress
press = AdaKVPress(
press=SnapKVPress(compression_ratio=0.5),
alpha_safeguard=0.20, # min 20% kept per head
)
# Note: does NOT reduce peak memory (uses fake keys)
# Note: requires NOT attn_implementation="eager"
StreamingLLM matching the original paper
from kvpress import StreamingLLMPress, KeyRerotationPress
press = KeyRerotationPress(
press=StreamingLLMPress(compression_ratio=0.8)
)
Per-layer compression ratios
from kvpress import PerLayerCompressionPress, SnapKVPress
ratios = [0.2] * 8 + [0.5] * 16 + [0.8] * 8 # 32 layers
press = PerLayerCompressionPress(
press=SnapKVPress(compression_ratio=0.0), # ratio overridden
compression_ratios=ratios,
)
Dynamic factory for multiple press types
from kvpress import (
ExpectedAttentionPress,
KnormPress,
SnapKVPress,
StreamingLLMPress,
TOVAPress,
)
PRESS_REGISTRY: dict[str, type] = {
"streaming_llm": StreamingLLMPress,
"snapkv": SnapKVPress,
"knorm": KnormPress,
"expected_attention": ExpectedAttentionPress,
"tova": TOVAPress,
}
def get_press(name: str, compression_ratio: float):
if name == "none":
return None
cls = PRESS_REGISTRY[name]
return cls(compression_ratio=compression_ratio)
Gotchas
-
compression_ratio is fraction to REMOVE.
0.9keeps only 10%. This is counterintuitive — double-check any code that sets this value. -
Prefill-only by default. The cache is compressed once during prefill. Tokens generated afterward all see the same compressed cache. If you need ongoing compression during generation, use
DecodingPress(experimental). -
ObservedAttentionPress needs eager attention. Load the model with
attn_implementation="eager"or it will assert-fail. This is significantly slower than flash/sdpa attention. -
AdaKVPress does NOT save memory. It uses "fake keys" (where
exp(<q,k>) ≈ 0) instead of actually removing entries. The cache stays the same size. It improves quality but not memory. -
ComposedPress limitations. Cannot contain AdaKVPress or KVzipPress. Presses that depend on attention weights may break if a prior press changes keys/values.
-
Model architecture requirement. The model must expose
model.model.layers[].self_attn. This is standard for Llama/Mistral/Qwen/Phi3 but not universal. -
Hooks persist until context exit. If an exception occurs inside
with press(model):, hooks are still cleaned up (finally block). But if you create hooks manually without the context manager, you must remove them yourself. -
Quantized caches work. kvpress handles
QuantizedCachetransparently — dequantizes before scoring, re-quantizes after compression. -
Batch size. All presses support
batch_size >= 1. Score tensors are shaped(batch, num_kv_heads, seq_len). -
Multi-GPU. Supported via
acceleratedevice_map. Hooks register on actual model layers regardless of device placement.
Score-Based vs Non-Score-Based
ScorerPress subclasses implement score(module, hidden_states, keys, values, attentions, kwargs) returning (batch, num_kv_heads, seq_len). Higher score = more important = kept. Bottom-k scored tokens are pruned via topk.
Categories of scoring:
- Position-based: StreamingLLMPress (sinks + recent)
- Key geometry: KnormPress, KeyDiffPress, LeverageScorePress, CURPress
- Attention-based: SnapKVPress, TOVAPress, ObservedAttentionPress, NonCausalAttnPress
- Statistical modeling: ExpectedAttentionPress
- Learned: QFilterPress, KVzapPress
- Random: RandomPress
Non-score presses use fundamentally different mechanisms: dimension pruning (ThinKPress), layer selection (SimLayerKVPress), head classification (DuoAttentionPress), multi-pass reconstruction (KVzipPress), or learned gates (FastKVzipPress).
Recommended Agent Skills
Expand your agent's capabilities with these related and highly-rated skills.
sparse-retrieval-eval
Evaluate sparse retrieval models on standard IR benchmarks (BEIR, MIRACL, mMARCO). Covers all IR metrics (nDCG@k, Recall@k, MAP, MRR), dataset loading, sparse corpus encoding to CSR matrices, IDF-weighted retrieval, caching, and result interpretation. Triggers on: evaluate retrieval, BEIR benchmark, nDCG, recall@k, sparse retrieval evaluation, MIRACL evaluation, information retrieval metrics, IR evaluation, search quality metrics.
wandb-plot
Download and generate plots from Weights & Biases runs. Use when you need to: - List projects you have access to - List runs in a W&B project - Inspect available metrics for a run - Download existing plot images from a run - Generate line plots from metric history (loss, accuracy, etc.)
go
Go engineering best practices and idioms. This skill should be used when writing, reviewing, or refactoring Go code. Triggers on any .go file work, Go module operations, go test, go build, go run, go vet, go generate, or when the user mentions Go, golang, goroutines, channels, context, Go interfaces, Go error handling, Go testing, or Go concurrency. ALWAYS use this skill when Go code is involved, even for simple functions.
fastapi
FastAPI best practices and conventions. Use when working with FastAPI APIs and Pydantic models for them. Keeps FastAPI code clean and up to date with the latest features and patterns, updated with new versions. Write new code or refactor and update old code.
plantuml
Create, edit, and render PlantUML diagrams. Triggers on: architecture diagrams, flowcharts, sequence diagrams, data models, state machines, visual documentation.
qdrant-sparse
Qdrant sparse vector operations: collection creation with SparseVectorParams, Modifier.IDF for miniCOIL/SPLADE/BM42, upserting SparseVector points, sparse search, hybrid search with prefetch + RRF/DBSF fusion, converting model outputs to SparseVector format, payload filtering, and performance tuning. Covers the sparse vector gap not handled by the official Qdrant MCP (which only supports dense vectors via FastEmbed).
Didn't find tool you were looking for?