Benchmarks · Language & multilingual

FLORES-200

also: Flores-200, No Language Left Behind evaluation benchmark

How well a translation system converts text between any pair of 200 languages, using the same set of professionally translated sentences in every language.

Meta AI (NLLB Team)Released 6 July 2022Live

Most machine-translation benchmarks before FLORES-200 covered a handful of high-resource language pairs, leaving the majority of the world’s languages with no standard way to measure translation quality at all. Meta AI built FLORES-200 to close that gap: around 2,000 sentences, translated by professional translators into 200 languages and dialects, aligned sentence-for-sentence so any language can be evaluated against any other — more than 40,000 possible directions from a single dataset.

The set was released alongside NLLB-200, a single model Meta trained to translate across all 200 languages, and the two were designed as a pair: FLORES-200 was the yardstick Meta used to report NLLB-200’s own results, claiming an average gain of 7.3 spBLEU points over the strongest prior systems, with improvements exceeding 70% for some African and Indian languages that had previously had little dedicated translation work done on them. FLORES-200 built directly on FLORES-101, an earlier 101-language version from the same Meta research group published the year before.

Because it was free, professionally translated and covered so many low-resource pairs at once, FLORES-200 became a standard reference set well beyond NLLB itself — researchers evaluating any multilingual model’s translation ability, including general-purpose language models rather than dedicated translation systems, routinely report chrF++ or spBLEU scores against it. Its sentences also became raw material for other multilingual benchmarks, including Meta’s own Belebele reading-comprehension set, which builds its passages directly from FLORES-200 text.

The set

Roughly 2,000 sentences drawn from Wikimedia sources, translated by professional translators into 200 languages and dialects, enabling many-to-many evaluation across more than 40,000 translation directions. Scored primarily with chrF++ and spBLEU.

Example

The first dev-set sentence, carried over unchanged when FLORES-101 was extended into FLORES-200: 'On Monday, scientists from the Stanford University School of Medicine announced the invention of a new diagnostic tool that can sort cells by type: a tiny printable chip that can be manufactured using standard inkjet printers for possibly about one U.S. cent each.', paired with its professional Icelandic translation, 'Á mánudag tilkynntu vísindamenn frá læknadeild Stanford-háskóla uppfinningu á nýju greiningartæki sem getur flokkað frumur eftir tegund: örlítill prentanleg flaga sem hægt er að framleiða með venjulegum bleksprautuprentara fyrir mögulega um eitt bandarískt sent stykkið.'dl.fbaipublicfiles.com

Where it stands

Still the default translation benchmark for low-resource and many-to-many evaluation; a 2026 paper flagged cross-direction contamination risk in some downstream uses.

How the top score changed hands

  1. July 2022NLLB-200+7.3 spBLEU average vs prior state of the artMeta reported gains of over 70% for some African and Indian languages against the strongest prior systems, evaluated on the FLORES-200 set built for the release.

In the timeline · 1 entry

More language & multilingual benchmarks