Anthropic open-sources its political even-handedness evaluation
Anthropic's own grading method scored Claude Sonnet 4.5 at 94% even-handedness, behind Gemini 2.5 Pro and Grok 4 but ahead of GPT-5 and Llama 4.
- Safety & alignment
- Colour
Anthropic published and open-sourced the methodology and dataset behind its “Paired Prompts” evaluation, which tests whether Claude engages with opposing political viewpoints at comparable depth, acknowledges counterarguments, and avoids refusing to answer.
Using Claude Sonnet 4.5 as an automated grader — cross-checked against GPT-5 and Claude Opus 4.1 — Anthropic reported Claude Sonnet 4.5 scored 94% on even-handedness, behind Gemini 2.5 Pro (97%) and Grok 4 (96%) but ahead of GPT-5 (89%) and Llama 4 (66%), with low refusal rates. Because Anthropic designed, ran and published the evaluation itself, the comparison is self-reported rather than independently verified. The company said it released the method to establish a shared industry benchmark for measuring political bias in AI systems, an area with no established external standard.