Model
Ant Ling releases Ling-3.0-flash, vLLM support coming soon
Ant Ling released Ling-3.0-flash, a 124B-parameter MoE model with only 5.1B activated parameters per token, featuring hybrid linear attention and native 256K context. The vLLM team confirmed that open-source support will be available simultaneously when the model weights are open-sourced, and praised Ant Ling's "announce first, open-source later" release model for providing the community with a stable adaptation window.
Read the original (opens in a new tab)
News stream data aggregated by AI HOT