Qwen3.8 is the August 2026 generation of 📝Qwen, Alibaba's large language model family, spanning the 2.4-trillion-parameter 📝Qwen3.8-Max flagship and a 27-billion-parameter dense open-weight model.
Built on the Qwen3.5 foundation, Qwen3.8 pairs a sparse 📝Mixture of Experts (MoE) design with hybrid attention — described in the model repository as Gated Delta Networks combined with sparse MoE — to hold inference cost down as total parameter counts climb. The Qwen team frames the generation around gains in coding, professional work, research, and long-horizon agentic tasks, with flexible reasoning controls.
Qwen3.8-Max was previewed at the World AI Conference in Shanghai on July 19, 2026 and launched as a hosted API on August 3. Open weights followed in two drops: Qwen3.8-2.4T-A95B on August 12, a text-only checkpoint with a 262,144-token native context rather than the API's one-million-token window, and Qwen3.8-27B on August 14, a dense model with a vision encoder that runs on a single GPU under the Apache 2.0 license. Qwen3.8-Flash variants are also listed as part of the generation.
Weights are distributed through Hugging Face and ModelScope with serving support in vLLM and SGLang, while hosted access through Alibaba Cloud Model Studio exposes endpoints compatible with both the 📝OpenAI and Anthropic API specifications, so the models slot into agent harnesses built for either. Coverage at launch widely framed the 27B checkpoint as the realistic on-premise deployment path, which places Qwen3.8 squarely inside the 📝Open-Weight vs. Closed-Weight AI Models tradeoff: the largest open checkpoint trails the API flagship on modality and context length.
