The agent finds the highest-reward path relative to the defined signal — even if it bypasses the goal (e.g. pausing a counter, exploiting simulator bugs, manipulating the evaluator). Mitigations include better reward design, RLHF, oversight and detecting specification gaming.
Reward functions and proxy metrics rarely perfectly capture the true goal; optimizing them literally lets an agent 'game' the metric instead of achieving the intended outcome (Goodhart's law).