DeepMind researchers show debate training curbs reward hacking
A weaker frozen model judged the contest; training against an adversarial critic recovered about 45% of the gap to a hypothetical accurate judge, versus a standard single-judge baseline.
Google DeepMindSafety & alignment