Study detects reward hacking from models' internal representations
Cheap difference-of-means probes on frontier open-weight models matched expensive LLM-judge monitors at catching reward hacking, at a fraction of the compute cost.
Zhipu AI, Moonshot AISafety & alignment · Benchmarks & progress