Do Transformers Need Three Projections? Redefining Attention Efficiency
Do Transformers Need Three Projections? Systematic Study of QKV Variants
This paper systematically investigates whether the standard tripartite Query, Key, and Value (QKV) projections are necessary for Transformers. The authors evaluate variants that share weights between projections, identifying that the Q-K=V (separate Query, shared Key-Value) configuration achieves a 50% KV cache reduction with parity or minimal performance degradation (e.g., +2.48% perplexity at 1.2B scale).
TL;DR
The classic tripartite architecture is not a law of nature. New research demonstrates that by forcing Key = Value (Q-K=V), we can slash the inference memory (KV cache) by 50% with almost no loss in model quality. When stacked with existing methods like Multi-Query Attention (MQA), memory savings reach a staggering 97%, effectively paving the way for running billion-parameter LLMs on local edge devices without the traditional "memory wall."
The Hidden Tax: The Complexity of QKV
Since the "Attention is All You Need" era, we have accepted that every token needs three linear projections: a Query () to ask, a Key () to be matched, and a Value () to provide content.
However, this creates a massive KV Cache bottleneck. During generation, the model must store every and vector for every previous token. For a 128k context window, this cache can consume over 10GB of VRAM for a relatively small model. The author's insight is simple: Is different enough from to justify its own separate memory footprint?
Methodology: Challenging the Trinity
The researchers systematically broke the QKV structure into three shared variants:
- Q=K-V: Shared Query and Key (Symmetric Attention).
- Q-K=V: Shared Key and Value (Asymmetric Attention).
- Q=K=V: One projection to rule them all.

Why Q-K=V Wins
The study found that Q-K=V is the "sweet spot." Unlike the version which forces the attention map to be symmetric (breaking the causal "flow" of language), allows the Query to remain independent. This preserves the directed nature of attention.
Empirically, the model learns to use the same representational space for both "addressing" (Keys) and "content" (Values). Analysis shows and in standard models already share high cosine similarity (~0.73), making them ideal candidates for weight tying.
Experimental Showdown: From Vision to 1.2B LLMs
The authors didn't just test on toy tasks. They scaled to 1.2 Billion parameters trained on 10 Billion tokens.
- Language Modeling: Q-K=V achieved 50% cache reduction with only a 2.48% increase in perplexity.
- Vision Tasks: On TinyImageNet, the simplest variant actually outperformed the standard baseline, suggesting vision tasks are even more redundant than language.
- Practical Speedup: On an A100 GPU, the Q-K=V variant increased decode throughput by ~5% and reduced total peak memory significantly.

The Synergy: Projection Sharing + Head Sharing
The most "Senior Architect" move in this paper is the demonstration of orthogonality. Modern LLMs use Grouped-Query Attention (GQA) to share heads. This paper shows you can do both.
| Method | KV Cache Reduction | Perplexity Change |
|---|---|---|
| Standard QKV | 0% | - |
| Q-K=V | 50% | +3.1% |
| MQA (Standard) | 93.8% | +1.5% |
| Q-MQA (Combined) | 96.9% | +4.8% |
At 1.2B scale, the combination of Q-MQA reduces the KV cache from 5.9 GB down to just 88 MB. This is the difference between needing a cluster of GPUs and running a model on a high-end smartphone.
Critical Analysis & Takeaways
The "Asymmetry" Insight
The primary takeaway is that Transformer efficiency isn't just about "fewer parameters"; it's about asymmetry. The must stay separate to handle the "search" logic, but the "index" () and the "data" () are close enough in high-dimensional space that they can be unified without the model collapsing.
Limitations
- Scale: While 1.2B is impressive for a systematic study, we have yet to see if the 2.4% perplexity gap shrinks or grows at the 70B+ scale.
- Linear Attention: The paper provides a fascinating theoretical link in the Appendix, showing that under QKV collapse, linear attention becomes a State-Space Model (SSM) with an adaptive readout.
Conclusion
This work provides a principled framework for trading a tiny bit of perplexity for a massive gain in deployability. For anyone building "Edge AI" or "On-device LLMs," Projection Sharing is likely the next standard optimization to enter the stack.
