Beyond µP: The Hidden Power of Embedding Layer Learning Rate

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate

2026-05-01
Dayal Singh Kalra, Maissam Barkeshli
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a quantitative framework to evaluate hyperparameter transfer in Large Language Models (LLMs) and identifies that the primary advantage of Maximal Update Parameterization (µP) over Standard Parameterization (SP) under AdamW stems almost entirely from the learning rate of the embedding layer. The authors demonstrate that scaling the embedding LR by a factor of width in SP effectively matches the transfer quality of µP.

TL;DR

Hyperparameter transfer is the "holy grail" of large-scale LLM training, allowing us to find optimal settings on tiny models and extrapolate them to trillions of parameters. While Maximal Update Parameterization (µP) is the industry standard for this, new research reveals a surprising secret: the vast majority of µP’s benefits under AdamW come solely from how it scales the embedding layer learning rate. By fixing the embedding bottleneck in Standard Parameterization (SP), we can achieve nearly identical transfer quality without the complexity of full µP.

The Bottleneck in the Foundation

Training massive models is prohibitively expensive, which is why we rely on Scaling Laws. However, Standard Parameterization (SP) often fails to transfer learning rates (LR) across widths because it suffers from "noisy" loss landscapes and instabilities at scale.

The authors argue that existing theories explaining why µP works—assumed to be about maintaining activation variance across all layers—are inadequate because they ignore practical realities like learning rate warmup and the compute-optimal (Chinchilla) regime.

A New Framework: The Three Pillars of Transfer

To move beyond qualitative "vibes," the paper introduces a rigorous metric system to judge any parameterization:

  1. Loss Predictability Error (E): Does the loss follow a smooth, predictable curve? (Lower is better).
  2. Transfer Robustness Exponent (κ): Does the loss landscape flatten (stable) or sharpen (brittle) as the model gets wider? (Negative κ is ideal).
  3. Asymptotic Loss Degradation (R(∞)): Does the parameterization actually reach the lowest possible loss at infinite scale?

The "Aha!" Moment: Isolating the Cultprit

The most striking part of this research is the systematic dismantling of µP. The authors compared SP and µP across four key differences:

  • Embedding LR Scaling
  • Last Layer Init Variance
  • LayerNorm LR Scaling
  • Attention Scaling (1/d vs 1/√d)

By testing all 16 possible combinations, they discovered that the Embedding Layer is the primary driver of success. In SP, the embedding LR typically scales as , which bottlenecks training. When this is changed to to match µP, SP suddenly becomes as stable and predictable as µP.

Table showing SP vs µP differences Table 1: The architectural differences between SP and µP. The authors identified the first row (Embedding) as the critical factor.

Visual Evidence: Fixing SP with One Layer

The charts below illustrate the transformation. Notice how standard SP has jagged, unpredictable curves. Once the embedding layer LR is "freed" (SP+Embd), the curves become smooth and align perfectly across widths—the hallmark of high-quality transfer.

Visual evidence of embedding LR importance Figure 2: Comparing SP, SP with fixed embeddings (SP+Embd), and µP. SP+Embd essentially recovers the stability of µP.

Why Does the Embedding Layer Matter So Much?

The embedding layer is a per-token lookup; unlike hidden layers, it doesn't involve a summation over the width . Therefore, scaling its LR by (as SP does) is "unnatural" and effectively freezes the layer's ability to learn meaningful representations in the early, critical phases of training. If the embedding isn't trained fast enough, it creates a "trash-in, trash-out" effect that destabilizes the entire downstream Transformer stack.

Critical Insight: The Chinchilla Conundrum

In the compute-optimal regime (scaling tokens in tandem with parameters), the authors found that current weight decay (WD) conventions are broken. As models get wider, the number of steps increases, leading to a "cumulative WD" effect that shifts the optimal learning rate. This suggests that as we move toward Chinchilla-optimal trillion-parameter models, even µP as we know it might need a redesign of its weight decay scaling ().

Conclusion & Application

This paper is a significant "de-mystifier" for LLM practitioners. The take-home message is simple:

  • If you use SP: Increase your embedding layer learning rate by a factor of width.
  • If you use µP: You're safe, but now you know why it's helping.
  • Future Research: The community needs to focus on weight decay scaling in the compute-optimal regime, as this remains the final frontier for perfect hyperparameter transfer.

Critical Note: While this holds for AdamW, different optimizers like SGD or Muon may have different "minimal variants" for stable transfer.

Find Similar Papers

Try Our Examples

  • Search for recent studies that explore how the embedding layer learning rate specifically affects the training stability and loss landscape of Transformer-based models.
  • Which paper originally proposed the Maximal Update Parameterization (µP) framework, and how does the current work's ablation study challenge its core theoretical justifications?
  • Investigate papers that propose alternative weight decay scaling laws for compute-optimal training regimes where the number of training steps scales with model width.
Contents
Beyond µP: The Hidden Power of Embedding Layer Learning Rate
1. TL;DR
2. The Bottleneck in the Foundation
3. A New Framework: The Three Pillars of Transfer
4. The "Aha!" Moment: Isolating the Cultprit
5. Visual Evidence: Fixing SP with One Layer
6. Why Does the Embedding Layer Matter So Much?
7. Critical Insight: The Chinchilla Conundrum
8. Conclusion & Application