Model
gemini-2.5-flash
Appears alongside
Featured in threads
Tracks
- Models & capabilities 2
- Safety & alignment 1
- Benchmarks & progress 1
DeepMind researchers show debate training curbs reward hacking
A weaker frozen model judged the contest; training against an adversarial critic recovered about 45% of the gap to a hypothetical accurate judge, versus a standard single-judge baseline.
Safety & alignment
ARC Prize compares reasoning models with no clear winner
ARC-AGI-2 remained unsolved by every system tested, and which model looked best depended entirely on whether accuracy or cost per task was prioritised.
Benchmarks & progress
Google I/O puts Gemini into search and ships Veo 3
AI Mode rolled out to all US Search users, and Veo 3 became the first widely-used video model to generate synchronised dialogue and sound effects alongside the picture.
Models & capabilities
Google launches Gemini 2.5 Flash in preview
A cheaper, faster sibling of Gemini 2.5 Pro with a configurable 'thinking budget' letting developers trade reasoning depth against cost and speed.
Models & capabilities