MIR & SoftQ: Solving the Data Drought in LLM Pretraining
Data-Constrained Language Model Pretraining: Improved Regularization and Scaling Laws
This paper investigates language model pretraining in data-constrained, compute-rich regimes, proposing Masked-Input Regularization (MIR) and the SoftQ scaling law. MIR improves autoregressive pretraining by adding an auxiliary masked-input loss, while SoftQ provides a superior fit for scaling behavior by coupling model and data size.
TL;DR
As we approach the limits of high-quality web data, LLM pretraining is shifting from "data-abundant" to "data-constrained." This paper by researchers at UMich and KAIST introduces Masked-Input Regularization (MIR)—a simple yet powerful auxiliary loss—and SoftQ, a new scaling law that finally accounts for the complex coupling between model size and finite data budgets.
Background: The End of the Single-Pass Era
For years, the industry followed the Chinchilla scaling law: if you have more compute, just get more data and a bigger model. But we are running out of unique human-generated text. Pretraining is now becoming compute-rich but data-constrained, forcing us to train for multiple epochs. Under these conditions, standard models suffer from "repeated-epoch collapse" or overfitting.
The authors identify two critical gaps:
- Regularization Gap: How do we stop models from memorizing specific training samples when they see them multiple times?
- Scaling Gap: Why do classical laws fail to predict performance when data is reused?
Methodology: MIR (Masked-Input Regularization)
The authors observe that Masked Diffusion Language Models (dLLMs) often resist overfitting better than Autoregressive (AR) models. They isolate the benefit of "masking" and port it to standard AR models via MIR.
The Core Intuition: In a data-constrained setting, a model might "memorize" a sequence using context-specific noise (like a specific typo or rare name) rather than learning the underlying logic. By masking parts of the input, MIR forces the model to predict the next token even when the "memorization cues" are missing.
The Math: The objective becomes: Where is the clean sequence and is the masked version.

Scaling Laws: Introducing SoftQ
The Chinchilla law is additive, meaning it assumes the benefit of data and model size are independent. However, the authors show a "fan-out" effect: the value of more data actually increases as models get bigger.
To fix this, they proposed SoftQ: This law introduces a bottleneck parameter that controls how model size () and unique data () interact.
Figure: SoftQ captures the data-model coupling that Chinchilla (dashed lines) misses.
Experimental Key Results
- Downstream Gains: At 1.4B parameters, MIR isn't just about lower loss; it translates to capability. It scored +10.2 points on BoolQ compared to a strongly regularized baseline.
- Data Efficiency: Using the SoftQ law, the authors calculated that MIR provides a gain equivalent to having 1.3x more unique training data.
- Token-Level Insight: Analysis shows MIR helps most in "noisy" contexts (mixed scripts, rare names, or technical text), where standard models typically fail by relying on simple frequency patterns.
Figure: MIR outperforms the baseline significantly on hard validation tokens.
Critical Insight & Conclusion
This work provides a theoretical and empirical bridge between Autoregressive and Masked modeling. By showing that masking is effectively a high-performance regularizer, it allows us to keep the efficient decoding of AR models while gaining the robustness of Diffusion models.
Takeaway: If you are training on a fixed dataset for multiple epochs, stop using Chinchilla coefficients and start masking your inputs. MIR is a "free lunch" in terms of data efficiency, provided you have the compute to spare for the extra forward pass.
