Model
gemini-2.5-flash-lite
Appears alongside
Featured in threads
Tracks
- Safety & alignment 1
- Models & capabilities 1
DeepMind researchers show debate training curbs reward hacking
A weaker frozen model judged the contest; training against an adversarial critic recovered about 45% of the gap to a hypothetical accurate judge, versus a standard single-judge baseline.
Safety & alignment
DeepMind's SIMA 2 uses Gemini to reason and act inside 3D game worlds
The agent, released as a limited research preview, generalised to games and AI-generated worlds it had not been trained on.
Models & capabilities