MIR & SoftQ: Solving the Data Drought in LLM Pretraining

Data-Constrained Language Model Pretraining: Improved Regularization and Scaling Laws

2026-06-01
Zhiwei Xu, Shihao Wu, Hanseul Cho, Wei Hu, Yixin Wang
Summary
Problem
Method
Results
Takeaways
Abstract

This paper investigates language model pretraining in data-constrained, compute-rich regimes, proposing Masked-Input Regularization (MIR) and the SoftQ scaling law. MIR improves autoregressive pretraining by adding an auxiliary masked-input loss, while SoftQ provides a superior fit for scaling behavior by coupling model and data size.

TL;DR

As we approach the limits of high-quality web data, LLM pretraining is shifting from "data-abundant" to "data-constrained." This paper by researchers at UMich and KAIST introduces Masked-Input Regularization (MIR)—a simple yet powerful auxiliary loss—and SoftQ, a new scaling law that finally accounts for the complex coupling between model size and finite data budgets.

Background: The End of the Single-Pass Era

For years, the industry followed the Chinchilla scaling law: if you have more compute, just get more data and a bigger model. But we are running out of unique human-generated text. Pretraining is now becoming compute-rich but data-constrained, forcing us to train for multiple epochs. Under these conditions, standard models suffer from "repeated-epoch collapse" or overfitting.

The authors identify two critical gaps:

  1. Regularization Gap: How do we stop models from memorizing specific training samples when they see them multiple times?
  2. Scaling Gap: Why do classical laws fail to predict performance when data is reused?

Methodology: MIR (Masked-Input Regularization)

The authors observe that Masked Diffusion Language Models (dLLMs) often resist overfitting better than Autoregressive (AR) models. They isolate the benefit of "masking" and port it to standard AR models via MIR.

The Core Intuition: In a data-constrained setting, a model might "memorize" a sequence using context-specific noise (like a specific typo or rare name) rather than learning the underlying logic. By masking parts of the input, MIR forces the model to predict the next token even when the "memorization cues" are missing.

The Math: The objective becomes: Where is the clean sequence and is the masked version.

Model Architecture and MIR Concept

Scaling Laws: Introducing SoftQ

The Chinchilla law is additive, meaning it assumes the benefit of data and model size are independent. However, the authors show a "fan-out" effect: the value of more data actually increases as models get bigger.

To fix this, they proposed SoftQ: This law introduces a bottleneck parameter that controls how model size () and unique data () interact.

Scaling Law Comparison Figure: SoftQ captures the data-model coupling that Chinchilla (dashed lines) misses.

Experimental Key Results

  • Downstream Gains: At 1.4B parameters, MIR isn't just about lower loss; it translates to capability. It scored +10.2 points on BoolQ compared to a strongly regularized baseline.
  • Data Efficiency: Using the SoftQ law, the authors calculated that MIR provides a gain equivalent to having 1.3x more unique training data.
  • Token-Level Insight: Analysis shows MIR helps most in "noisy" contexts (mixed scripts, rare names, or technical text), where standard models typically fail by relying on simple frequency patterns.

MIR Token-Level Performance Figure: MIR outperforms the baseline significantly on hard validation tokens.

Critical Insight & Conclusion

This work provides a theoretical and empirical bridge between Autoregressive and Masked modeling. By showing that masking is effectively a high-performance regularizer, it allows us to keep the efficient decoding of AR models while gaining the robustness of Diffusion models.

Takeaway: If you are training on a fixed dataset for multiple epochs, stop using Chinchilla coefficients and start masking your inputs. MIR is a "free lunch" in terms of data efficiency, provided you have the compute to spare for the extra forward pass.

Find Similar Papers

Try Our Examples

  • Examine recent literature on "data-constrained" vs "compute-optimal" pretraining strategies specifically for Frontier-scale models.
  • What are the foundational theories behind the "Quanta" hypothesis in neural scaling laws, and how does the SoftQ law mathematically extend them?
  • Search for studies applying dual masked-autoregressive objectives in domains like computer vision or reinforcement learning to mitigate overfitting on small datasets.
Contents
MIR & SoftQ: Solving the Data Drought in LLM Pretraining
1. TL;DR
2. Background: The End of the Single-Pass Era
3. Methodology: MIR (Masked-Input Regularization)
4. Scaling Laws: Introducing SoftQ
5. Experimental Key Results
6. Critical Insight & Conclusion