Do Transformers Need Three Projections? Redefining Attention Efficiency

Do Transformers Need Three Projections? Systematic Study of QKV Variants

2026-06-01
Ali Kayyam, Anusha Madan Gopal, M Anthony Lewis
Summary
Problem
Method
Results
Takeaways
Abstract

This paper systematically investigates whether the standard tripartite Query, Key, and Value (QKV) projections are necessary for Transformers. The authors evaluate variants that share weights between projections, identifying that the Q-K=V (separate Query, shared Key-Value) configuration achieves a 50% KV cache reduction with parity or minimal performance degradation (e.g., +2.48% perplexity at 1.2B scale).

TL;DR

The classic tripartite architecture is not a law of nature. New research demonstrates that by forcing Key = Value (Q-K=V), we can slash the inference memory (KV cache) by 50% with almost no loss in model quality. When stacked with existing methods like Multi-Query Attention (MQA), memory savings reach a staggering 97%, effectively paving the way for running billion-parameter LLMs on local edge devices without the traditional "memory wall."

The Hidden Tax: The Complexity of QKV

Since the "Attention is All You Need" era, we have accepted that every token needs three linear projections: a Query () to ask, a Key () to be matched, and a Value () to provide content.

However, this creates a massive KV Cache bottleneck. During generation, the model must store every and vector for every previous token. For a 128k context window, this cache can consume over 10GB of VRAM for a relatively small model. The author's insight is simple: Is different enough from to justify its own separate memory footprint?

Methodology: Challenging the Trinity

The researchers systematically broke the QKV structure into three shared variants:

  1. Q=K-V: Shared Query and Key (Symmetric Attention).
  2. Q-K=V: Shared Key and Value (Asymmetric Attention).
  3. Q=K=V: One projection to rule them all.

Projection Variants

Why Q-K=V Wins

The study found that Q-K=V is the "sweet spot." Unlike the version which forces the attention map to be symmetric (breaking the causal "flow" of language), allows the Query to remain independent. This preserves the directed nature of attention.

Empirically, the model learns to use the same representational space for both "addressing" (Keys) and "content" (Values). Analysis shows and in standard models already share high cosine similarity (~0.73), making them ideal candidates for weight tying.

Experimental Showdown: From Vision to 1.2B LLMs

The authors didn't just test on toy tasks. They scaled to 1.2 Billion parameters trained on 10 Billion tokens.

  • Language Modeling: Q-K=V achieved 50% cache reduction with only a 2.48% increase in perplexity.
  • Vision Tasks: On TinyImageNet, the simplest variant actually outperformed the standard baseline, suggesting vision tasks are even more redundant than language.
  • Practical Speedup: On an A100 GPU, the Q-K=V variant increased decode throughput by ~5% and reduced total peak memory significantly.

Efficiency-Quality Pareto Frontier

The Synergy: Projection Sharing + Head Sharing

The most "Senior Architect" move in this paper is the demonstration of orthogonality. Modern LLMs use Grouped-Query Attention (GQA) to share heads. This paper shows you can do both.

MethodKV Cache ReductionPerplexity Change
Standard QKV0%-
Q-K=V50%+3.1%
MQA (Standard)93.8%+1.5%
Q-MQA (Combined)96.9%+4.8%

At 1.2B scale, the combination of Q-MQA reduces the KV cache from 5.9 GB down to just 88 MB. This is the difference between needing a cluster of GPUs and running a model on a high-end smartphone.

Critical Analysis & Takeaways

The "Asymmetry" Insight

The primary takeaway is that Transformer efficiency isn't just about "fewer parameters"; it's about asymmetry. The must stay separate to handle the "search" logic, but the "index" () and the "data" () are close enough in high-dimensional space that they can be unified without the model collapsing.

Limitations

  • Scale: While 1.2B is impressive for a systematic study, we have yet to see if the 2.4% perplexity gap shrinks or grows at the 70B+ scale.
  • Linear Attention: The paper provides a fascinating theoretical link in the Appendix, showing that under QKV collapse, linear attention becomes a State-Space Model (SSM) with an adaptive readout.

Conclusion

This work provides a principled framework for trading a tiny bit of perplexity for a massive gain in deployability. For anyone building "Edge AI" or "On-device LLMs," Projection Sharing is likely the next standard optimization to enter the stack.

Find Similar Papers

Try Our Examples

  • Search for recent papers that investigate "weight tying" or "parameter sharing" specifically within the attention mechanism projections beyond the QKV setup.
  • Which paper first formally introduced Multi-Query Attention (MQA) and how does its architectural constraint on the KV cache compare to the hard equality constraint of Q-K=V?
  • Explore if current state-of-the-art long-context models like Gemini or GPT-4 utilize projection-sharing techniques to mitigate KV cache growth.
Contents
Do Transformers Need Three Projections? Redefining Attention Efficiency
1. TL;DR
2. The Hidden Tax: The Complexity of QKV
3. Methodology: Challenging the Trinity
3.1. Why Q-K=V Wins
4. Experimental Showdown: From Vision to 1.2B LLMs
5. The Synergy: Projection Sharing + Head Sharing
6. Critical Analysis & Takeaways
6.1. The "Asymmetry" Insight
6.2. Limitations
7. Conclusion