Techniques include reinforcement learning from human feedback (RLHF) and from AI feedback (RLAIF), constitutional AI, scalable oversight, interpretability, red-teaming, and capability/threat evaluations (e.g. work by METR and Palisade Research).
It is hard to precisely encode human intentions and values into an optimization objective; misaligned, capable systems may act in undesirable or unsafe ways.