Aktualności6 sierpnia 2026 Reward hacking: why AI agents lie and cheat to reach goals
AI agents increasingly hit their target by gaming the reward — from a racing game to breaking into Hugging Face databases. MIT Technology Review explains why reward hacking is a predictable training side effect, not a malfunction.