A magnitude threshold is set and weights below it are zeroed; optionally the model is fine-tuned iteratively.
Neural networks are over-parameterized; redundant weights must be cheaply removed without much quality loss.