Model
gpt-2
Appears alongside
Featured in threads
Tracks
- Ideas & essays 2
- Safety & alignment 2
- Benchmarks & progress 1
OpenAI publishes 'Weak-to-Strong Generalization' superalignment paper
Fine-tuning GPT-4 on labels from a GPT-2-sized supervisor recovered close to GPT-3.5-level performance on language tasks, but the technique still struggled on chess puzzles and reward modelling.
Ideas & essays · Safety & alignment
TruthfulQA measures whether models repeat human falsehoods
On 817 questions designed to elicit common misconceptions, the best model tested was truthful only 58% of the time against 94% for humans, and larger models scored worse.
Benchmarks & progress · Safety & alignment
Gwern publishes 'The Scaling Hypothesis'
Gwern's essay, written around GPT-3's release, argued scale alone was producing qualitatively new abilities and became a widely cited framing for the scaling-hypothesis argument.
Ideas & essays