Timeline

Baidu announces ERNIE 3.0 Titan

Built with Peng Cheng Laboratory, the 260-billion-parameter model reported state-of-the-art results on more than 60 Chinese-language NLP tasks.

  • Models & capabilities
  • Minor

Baidu, working with the state-backed Peng Cheng Laboratory, published a paper describing ERNIE 3.0 Titan, a 260-billion-parameter Chinese-language model the authors called the largest dense pre-trained model built in China at the time. It extended an earlier, roughly 10-billion-parameter version of ERNIE 3.0 released the same year, and the authors reported state-of-the-art results across more than 60 Chinese natural-language-processing tasks, including reading comprehension and text classification.

The paper’s technical contributions went beyond scale. The authors introduced a self-supervised adversarial training method intended to help the model distinguish genuine human-written text from machine-generated text, and an online distillation framework that trained smaller “student” models alongside the large “teacher” model rather than after it, which they said cut the computational cost of producing deployable smaller versions. A distilled student model was reported to outperform BERT-Base and RoBERTa-Base by several percentage points on some tasks despite having far fewer parameters.

The release positioned Baidu as the most visible Chinese counterpart to OpenAI and Google in building GPT-3-class models, at a time when Chinese labs — including the Beijing Academy of Artificial Intelligence’s Wu Dao project earlier the same year — were demonstrating that frontier-scale language modelling was not confined to a small number of US labs. ERNIE 3.0 Titan’s parameter count, while large by the standards of late 2021, was soon superseded on both sides of the Pacific as the industry’s scaling race continued into 2022.