Moonshot AI releases Kimi Linear architecture model
Moonshot's hybrid attention design cut KV-cache memory by up to 75% and lifted decoding speed up to sixfold at 1-million-token context, released with open weights and kernels.
Open weights & ecosystem · Models & capabilities