Robots Atlas>ROBOTS ATLAS
Artificial Intelligence

DeepSeek-V4-Flash really uses two of four mHC streams

Sir Robot3 October 2026 · 3 min read
DeepSeek-V4-Flash really uses two of four mHC streams

Huawei engineers measured how DeepSeek-V4-Flash actually uses its four-stream mHC residual pathway. A typical block effectively uses about two streams, and the residual mixers in layers 22–42 can be swapped for an identity matrix with no loss of quality. mHC is one of three pillars of the DeepSeek-V4 architecture.

Key takeaways

  • Huawei Technologies analysis posted to arXiv on 4 September, using the public 0731 DeepSeek-V4-Flash checkpoint
  • Model: 43-layer MoE, 284B parameters (13B activated), 86 instrumented sites
  • Load-effective stream count: 1.998 read, 1.775 write, out of four available
  • Layer 22–42 mixers replaced by identity: six-task average rises from 84.22 to 84.26 points
  • The same swap in layers 0–21 raises C4 perplexity by 41.4 percent, cutting the average to 80.94

Four streams, two actually loaded

mHC, or manifold-constrained Hyper-Connections, widens a transformer's single residual stream into four parallel ones. The measured load spreads across just 1.998 read streams and 1.775 write streams.

1.998load-effective read streams out of four available — 1.775 on writearXiv:2609.05309

This is not one globally dominant stream. Streams 0 and 1 dominate 32 of 44 sites in layers 0–21, streams 2 and 3 take over 36 of 42 sites later, and a 0.404 cosine similarity keeps the representations directionally distinct. An earlier study of 120M and 360M models reported a hard collapse onto one stream. At 284B this is uneven loading, not collapse.

Late mixing is redundant, early mixing is not

Mixer deviation from Identity matrix: A square matrix with ones on the diagonal and zeros elsewhere. Multiplying by it changes nothing, so each stream passes through without mixing with the others. falls from 0.046 in layers 0–21 to 0.009 in layers 22–42. Replacing the late mixers with identity raises C4 perplexity by 1.9 percent, while the six-benchmark average (ARC-Easy, ARC-Challenge, PIQA, HellaSwag, MMLU, GSM8K) rises from 84.22 to 84.26. The same swap early costs 41.4 percent perplexity and 3.28 average points. In practice: the model's second half can skip Sinkhorn–Knopp iterations at inference.

…
Symbol meaning
…
residual mixer of layer l — the matrix that mixes the four streams
…
the 4×4 identity matrix, meaning no mixing: each stream carries on separately
…
layer index, the model has 43 layers numbered 0 to 42
VariantSix-task averageC4 perplexity
Baseline — mHC unchanged84.22reference
Layer 22–42 mixers → identity matrix84.26+1.9 percent

Why it matters

Architectural width does not automatically become usable capacity. How many of the four streams work is settled by training, not design. For teams designing the next models, some of mHC's expensive degrees of freedom can be cut with no quality loss and the saving collected at inference. An architectural result from a launch paper says little about how the mechanism behaves after 32 trillion training tokens.

What's next

  • Huawei proposes conditioning the write map, and possibly the mixer, on the branch output rather than only on incoming states — this needs matched training runs.
  • The field is diverging: Hy4-preview uses Hyper-Connections with an identity mixer, Qwen3.8-Next a data-dependent gated residual.
  • The analysis covers one checkpoint, so telling mHC's properties apart from artefacts of this run needs training-time data.

Sources

Share this article