Shanghai AI Laboratory releases InternLM3
An 8B open-weight model trained on 4 trillion tokens that Shanghai AI Lab said matched rivals trained on far more data, cutting training cost by over 75%.
- Open weights & ecosystem
- Models & capabilities
- Colour
Shanghai AI Laboratory released InternLM3-8B-Instruct, an open-weight 8-billion-parameter model continuing its InternLM line and combining conversational and explicit “deep-thinking” reasoning modes in a single checkpoint, released under Apache 2.0.
The lab’s headline claim was efficiency rather than raw capability: trained on 4 trillion tokens, it said InternLM3 matched the performance of models trained on roughly 18 trillion tokens elsewhere in the field, and cut training cost by more than 75% against comparable open models. It reported a data-efficiency measure more than four times higher than Meta’s Llama 3.1 and put overall performance close to GPT-4o-mini across the benchmarks it tested.
The release was one of several January 2025 open-weight models from Chinese labs — alongside DeepSeek’s and Alibaba’s releases the same month — that argued their case on cost-per-token and training efficiency rather than outright leaderboard position, part of a broader shift in which data efficiency and inference cost became competitive axes alongside benchmark scores.