From One to Millions: Scaling Personal Models on Trillion-Parameter Priors
On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters
This paper introduces a paradigm shift in PEFT, positioning small adapters as persistent local state for personal models rather than just cheap fine-tuning alternatives. Through the Mind Lab's "MinT" infrastructure, they demonstrate "Scale Up, Scale Down, and Scale Out" strategies, achieving SOTA results in trillion-scale MoE reinforcement learning and million-user simulation.
TL;DR
The frontier of AI is shifting from building "one model to rule them all" to "millions of models that know you." This paper from Mind Lab presents a comprehensive blueprint for Parameter-Efficient Fine-Tuning (PEFT) as the cornerstone of persistent personal identity. By scaling Up (1T+ MoE models), Down (rank-1 stability), and Out (addressing millions of users), they demonstrate how millions of unique "adapter-brains" can thrive on a single shared biological-scale prior.
Evolution of the Personal Model: Why Prompts Aren't Enough
A capable assistant is not automatically a personal one. While long contexts and RAG help, they are transient. To achieve true continuity, a model needs persistent state. The authors argue that LoRA (Low-Rank Adaptation) is not just a budget hack; it is the "genetic variation" (the 0.1% difference) that allows one shared "human biology" (the base model) to support billions of individuated lives.
The Three-Axis Scaling Framework
1. Scale Up: Bridging the Trillion-Parameter Gap
Personalization is high-leverage only if the base model is powerful. The authors successfully implemented LoRA RL on a 1.04 Trillion parameter MoE (Mixture-of-Experts).
- The Problem: Training-Inference Mismatch (TIM). In MoE, if the training engine routes tokens differently than the inference engine, the gradients become meaningless.
- The Solution: Router Replay (R3). By recording routing decisions during rollout and replaying them during training, they stabilized the KL divergence and ensured valid policy updates.

2. Scale Down: Stability at the Edge of Collapse
To support millions of users, adapters must be tiny (Rank 1). Usually, Rank 1 adapters are unstable and fail across different seeds.
- OLoRA-tail: The authors found that standard initialization wastes the limited "direction" of a Rank-1 adapter. By focusing on the minor singular vectors (the "tail" of the representation) and removing aggressive scaling, they made Rank-1 adaptation reliable even at large batch sizes.
- δ-mem: Beyond static weights, they introduced stateful adapters that act like a "working memory," writing residuals of the history into a compact associative state.

3. Scale Out: Diversity as Collective Intelligence
When you have 200 different adapters, you have more than just 200 users; you have a "Council of Experts."
- Collective Performance: By aggregating the majority vote of 198 diverse LoRA variants, the team saw a massive jump in AIME24 accuracy (from 36% to 48%). This proves that distinct adapter trajectories learn complementary reasoning paths that simple "repeated sampling" from a single model cannot replicate.
- User Simulation: In the OASIS social simulator, per-user LoRAs prevented the "personality collapse" seen in prompt-only agents, producing richer interaction topologies and more realistic "echo chamber" dynamics.

Infrastructure: The MinT System
You cannot manage a million personal models with a folder full of files. Mind Lab's MinT infrastructure treats policies as identities with separate lifecycles:
- Resident Base Models: Keep the 1T+ model warm in memory.
- Tiered Residency: Orchestrate adapters between GPU slots (hot), CPU cache (warm), and shared storage (cold).
- Two-Phase Readiness: Pre-warming adapters before exposing them to users to prevent "TTFT stalls" that ruin user experience.
Critical Insight & Future Outlook
The core takeaway is that Diversity is a Resource. By making adaptation cheap (Scale Down) and infrastructure robust (MinT), we stop trying to build a single "perfect" model and instead foster a population of specialized ones.
Future Challenges:
- Memory Capacity: LoRA has a "hard ceiling" for how many facts it can store (roughly tokens per parameter). We need better "Write Policies" (Context Learning) to decide what deserves to be carved into the weights.
- Drift: Ensuring that millions of personal models don't eventually regress to the "average" prior over long interaction horizons.
This paper provides the technical foundation for a future where your AI isn't just a leased endpoint from a big tech company, but a persistent, evolving extension of your own digital identity.
