Qwen has released a 125B-parameter model that activates just 6B. The company calls it an early preview of the Qwen4 architecture.