DeepSeek open-sources its first native vision model
Weights followed the API by ten days; DeepSeek shipped a reference PyTorch inference implementation alongside the MIT-licensed checkpoint.
- Open weights & ecosystem
- Models & capabilities
- Minor
DeepSeek published open weights for DeepSeek-V4-Flash-Vision-Exp, which the company described as its first experimental multimodal release in the V4 family, adding a vision encoder to the existing V4-Flash model rather than shipping visual understanding as a separate pipeline. The checkpoint was released under an unmodified MIT licence, alongside a reference PyTorch inference implementation covering the vision encoder and generation logic — described in the repository as readable code rather than a production serving engine.
The release was staged: DeepSeek made the model available through its API on 21 August, saying it matched V4-Flash on text tasks such as reasoning and agentic tool use while bringing multimodal agent performance closer to rival closed models, before publishing the underlying weights ten days later. DeepSeek reported gains on its own multimodal agent benchmarks compared with the prior V4-Flash checkpoint, including scores on ApexBench and a chart-reading evaluation; these are the company’s figures and have not been independently reproduced.
The release added DeepSeek to the list of Chinese labs shipping native multimodal capability under permissive open licences during 2026, continuing a run of frequent open-weight updates that had made individual DeepSeek releases increasingly routine.