For each weight a score |weight| × input-activation norm is computed; per output, the lowest-scoring weights are removed, with no retraining.
Classic magnitude-only pruning works poorly for LLMs, while second-order methods are expensive.