PAPER DEEP DIVE
Scaling Native Multimodal Pre-Training From Scratch
Although large language models (LLMs) exhibit remarkable reasoning capabilities, their reliance on text-only pre-training restricts the perception of the multimodal physical world. Native multimodal pre-training avoids this limitation by training models from scratch on multimodal inputs, thereby achieving deep cross-modal integration and mitigating optimization asymmetries inherent to traditional late-fusion architectures. Despite these advantages, the scaling properties of this paradigm remain systematically uncharacterized. To address this gap, we investigate the optimal model size and token count for training a transformer-based vision-language model under a fixed computational budget. We demonstrate that minimal objective loss adheres to a predictable compute law, whereas compute-optimal model sizes and token counts scale as power laws. Notably, language and multimodal objectives manifest distinct scaling behaviors. The language allocation law is largely invariant to the composition of the data, indicating stable language learning regardless of the multimodal data ratio. Conversely, the multimodal allocation law is highly sensitive to this composition. Specifically, text-heavy mixtures become compute-efficient only at larger model scales, shifting the optimal resource allocation toward greater model capacity. Additionally, by modeling the influence of data composition on compute laws and allocation exponents, we derive an efficiency frontier specifying precise configurations of model size, token count, and data mixture. Downstream evaluations further reveal that native multimodal pre-training induces positive cross-modal transfer, thereby enhancing pure-text spatial reasoning and enabling robust multimodal in-context learning. In summary, this empirical research establishes the essential groundwork for predictably scaling multimodal foundation models.
Scaling Native Multimodal Pre-Training From Scratch
Paper: Scaling Native Multimodal Pre-Training From Scratch
Authors: Haoyuan Wu, Aoqi Wu, Hai Wang, Jiajia Wu, Jinxiang Ou, Bei Yu
Affiliations: The Chinese University of Hong Kong / Tencent LLM Department
Links: arXiv:2607.22043
One-Sentence Summary
This paper systematically characterizes the scaling laws of native multimodal pre-training, finding that under a fixed compute budget the language objective's optimal allocation is a composition-invariant power law while the multimodal objective's is highly composition-sensitive — text-heavy mixtures become compute-efficient only at larger scales — and derives an efficiency frontier specifying precise configurations of model size, token count, and data mixture, while discovering that native multimodal pre-training induces positive cross-modal transfer enhancing pure-text spatial reasoning and enabling robust multimodal in-context learning.
Background and Motivation
Large language models exhibit remarkable reasoning capabilities, but text-only pre-training is fundamentally constrained by the inability to ground multimodal concepts in the physical world. The dominant multimodal pre-training paradigm relies on late fusion — coupling a pre-trained language model with a vision encoder via a projection layer, then continuing training on multimodal data. While efficiently leveraging existing weights, this introduces a fundamental asymmetry: vision and language representations are learned independently from distinct distributions under different objectives.
Native multimodal pre-training resolves this asymmetry by training from scratch, enabling deep modality integration and shared representational capacity. However, the scaling properties of this paradigm remain systematically uncharacterized. The core practical question: under a fixed compute budget, how to optimally allocate resources between model size and data (token count)? For unimodal language models, this is guided by compute-optimal scaling laws (Chinchilla etc.), but in native multimodal pre-training where language and multimodal objectives compete for shared parameters, the optimal allocation remains an open question.
This paper investigates the optimal model size and token count for a transformer-based vision-language model under a fixed compute budget, revealing distinct scaling behaviors for language and multimodal objectives, and deriving an efficiency frontier. Experiments cover model scales from 71M to 3B, data mixture ratios $r$ from 0 to 0.3, with up to 250B text and 75B multimodal tokens. Evaluation spans 16 text benchmarks and 21 multimodal benchmarks including spatial reasoning, coding, math, and logic.
Figure 1: IsoFLOP curves (language objective). Training tokens adjusted per model size to maintain constant final FLOPs.
Method
1. Compute-Optimal Allocation Framework
Let $N$ be model parameters, $D$ training tokens, and $r$ the data mixture ratio (multimodal token fraction). Total compute budget $C_{\text{total}}$ splits into text compute $C_{\text{text}}$ and multimodal compute $C_{\text{mm}}$. The compute-optimal scaling law takes the form:
$$ L(C) = \alpha \cdot C^{-\beta}, \quad N_{\text{opt}} \propto C^{a}, \quad D_{\text{opt}} \propto C^{b}, \quad a + b \approx 1 $$where $a$ is the parameter scaling exponent and $b$ the data scaling exponent. The minimal objective loss follows a predictable compute law:
$$ L_{\text{text}} = \alpha_t \cdot C_{\text{text}}^{-\beta_t}, \quad L_{\text{mm}} = \alpha_m(r) \cdot C_{\text{mm}}^{-\beta_m(r)} $$where the language objective's $\alpha_t, \beta_t$ are invariant to $r$, while the multimodal objective's $\alpha_m, \beta_m$ depend on $r$. The total loss is a weighted combination with weights determined by $r$.
2. Language Objective: Composition-Invariant
Figure 3: Compute-optimal allocation (language objective). (a) Optimal model size $N_{\text{opt}}$ and (b) optimal token count $D_{\text{opt}}$ as functions of $C_{\text{text}}$. Slopes show minimal variation with $r$.
For each model size $N$, training tokens $D$ are adjusted to maintain constant final FLOPs. The IsoFLOP curves show stable parabolic minima satisfying:
$$ \frac{\partial L}{\partial N}\bigg|_{N=N_{\text{opt}}} = 0, \quad N_{\text{opt}} = \arg\min_N L(N, D(N)) $$across all $r$ values. Key finding: the language objective's compute-optimal allocation is highly composition-invariant — the parameter capacity needed to minimize $L_{\text{text}}$ is strictly governed by the isolated budget $C_{\text{text}}$, showing empirical robustness to multimodal token introduction. Both estimation methods (IsoFLOP fits and training-curve envelope) confirm no monotonic downward trend.
3. Multimodal Objective: Composition-Sensitive
Figure 4: IsoFLOP curves (multimodal objective). The distinct valley demonstrates an optimal model size exists for any given FLOP budget.
In stark contrast, the multimodal objective's scaling behavior is strictly composition-variant. As $r$ increases from 0.1 to 0.3, the optimal model size exponent decreases substantially, validated by the training-curve envelope showing a consistent decline:
$$ a_{\text{mm}}(r): \quad r=0.1 \to \text{higher}, \quad r=0.3 \to \text{significantly lower} $$This reflects the inherently data-hungry nature of cross-modal alignment: processing increasingly dense multimodal data continuously flattens the optimal parameter-scaling curve, shifting optimal compute allocation from parameter expansion toward data expansion. Text-heavy mixtures become compute-efficient only at larger scales.
4. Joint Pareto Frontier
Figure 6: Compute-optimal allocation (multimodal objective). Slopes vary significantly with $r$, demonstrating composition sensitivity.
Practical pre-training must optimize shared parameters $N$ under unified budget $C_{\text{total}}$. Competition between $L_{\text{text}}$ and $L_{\text{mm}}$ establishes a Pareto frontier across $r$. Asymmetric modeling pairs the composition-invariant language objective with the composition-variant multimodal objective. Global optimal allocation exponents $a(r)$ and $b(r)$ reveal:
$$ r=0.1: \quad N_{\text{opt}} \propto C_{\text{total}}^{0.69}; \qquad r=0.3: \quad N_{\text{opt}} \propto C_{\text{total}}^{0.66}, \quad D_{\text{opt}} \propto C_{\text{total}}^{0.34} $$
flowchart TD
A["Fixed compute budget C_total"] --> B["Data mixture ratio r
multimodal token fraction"]
B --> C["Text compute C_text"]
B --> D["Multimodal compute C_mm"]
C --> E["Language objective L_text
composition-invariant scaling"]
D --> F["Multimodal objective L_mm
composition-sensitive scaling"]
E --> G["a_text invariant to r
capacity governed by C_text"]
F --> H["a_mm(r) decreases with r
data hunger flattens param curve"]
G --> I["Joint Pareto frontier
a(r) + b(r) ≈ 1"]
H --> I
I --> J["Efficiency frontier
precise N_opt, D_opt, r config"]
J --> K["r=0.1: N∝C^0.69
r=0.3: N∝C^0.66, D∝C^0.34"]
style E fill:#e1f5fe
style F fill:#fff3e0
style J fill:#e8f5e9
Experiments
Downstream Evaluation
Figure 2: Training curve envelope (language objective). Minimal loss per FLOP extracted to estimate optimal allocation.
| Objective | Composition Sensitivity | Param Exponent $a$ | Data Exponent $b$ | Key Finding |
|---|---|---|---|---|
| Language $L_{\text{text}}$ | invariant | invariant to $r$ | invariant to $r$ | capacity governed by $C_{\text{text}}$ |
| Multimodal $L_{\text{mm}}$ | sensitive | $r=0.1$: higher; $r=0.3$: lower | complementarily larger | cross-modal alignment is data-hungry |
| Joint ($r=0.1$) | - | 0.69 | 0.31 | parameter expansion favored |
| Joint ($r=0.3$) | - | 0.66 | 0.34 | data expansion more aggressive |
Cross-Modal Transfer
| Finding | Details |
|---|---|
| Text performance preserved | Models with different $r$ differ by <1pp on 16 text benchmarks, no interference |
| Spatial reasoning transfer | $r=0.3$ multimodal models consistently outperform $r=0$ baselines on SpatialEval (MazeNav, SpatialMap) text-only tasks, advantage grows with scale |
| ICL emergence | 3B model gains +2.43pp from 3-shot vs 0-shot; small models show no benefit |
| Few-shot concentrated in spatial | Spatial reasoning benchmarks improve most; OCR/recognition plateau or degrade |
Figure 5: Training curve envelope (multimodal objective) across varying $r$ and model sizes.
Limitations
- Model scale ceiling: Experiments max out at 3B parameters; the extrapolation accuracy of derived scaling laws at larger scales (e.g., 70B+) is unverified, and power-law extrapolation may carry偏差.
- Limited modality coverage: Only vision-language bimodal is studied; scaling behavior of native pre-training with audio, video, or other modalities is not covered.
- Simplified mixture dimension: Only the multimodal token ratio $r$ is used as the mixture parameter; quality and diversity differences within multimodal data are not considered.
Conclusion and Outlook
This paper systematically characterizes the scaling laws of native multimodal pre-training, revealing the fundamental asymmetry between the language objective's composition-invariance and the multimodal objective's composition-sensitivity. Through the joint Pareto frontier, it derives an efficiency frontier specifying precise model size, token count, and data mixture configurations. Downstream evaluation discovers positive cross-modal transfer from native multimodal pre-training — enhancing pure-text spatial reasoning and enabling multimodal in-context learning to emerge with scale. This establishes the groundwork for predictably scaling multimodal foundation models.
Golden quote: "The language objective's scaling is composition-invariant — parameter capacity is strictly governed by the text compute budget; while the multimodal objective's scaling is highly composition-sensitive — the data hunger of cross-modal alignment flattens the parameter curve as multimodal ratio increases, forcing optimal allocation to shift from parameter expansion to data expansion."
SOURCE LINKS

