{"id":648843,"date":"2026-08-26T17:39:50","date_gmt":"2026-08-26T17:39:50","guid":{"rendered":"https:\/\/Blockchain.News\/news\/alibaba-qwen3-8-flash-next-release"},"modified":"2026-08-26T17:39:50","modified_gmt":"2026-08-26T17:39:50","slug":"alibaba-unveils-qwen3-8-flash-next-ai-model-with-176b-parameters","status":"publish","type":"post","link":"https:\/\/e-bitco.in\/index.php\/2026\/08\/26\/alibaba-unveils-qwen3-8-flash-next-ai-model-with-176b-parameters\/","title":{"rendered":"Alibaba Unveils Qwen3.8-Flash-Next AI Model with 176B Parameters"},"content":{"rendered":"<figure class=\"figure mt-2\">\n<p> <a href=\"https:\/\/blockchain.news\/Profile\/Darius-Baruo\">Darius Baruo<\/a> <span class=\"publication-date ml-2\"> Aug 26, 2026 17:39<\/span> <\/p>\n<p class=\"lead\">Alibaba releases Qwen3.8-Flash-Next, a 176B-parameter Mixture-of-Experts model, previewing Qwen4 architecture and optimized for long-context tasks.<\/p>\n<p> <a href=\"https:\/\/image.blockchain.news\/features\/45B7A801D37E36DC0019AE0310A0ED0160FBF51AC1E55381847ABA1D7FFAC0B0.jpg\" class=\"hero-image-link\"> <img fetchpriority=\"high\" decoding=\"async\" class=\"rounded hero-image\" src=\"https:\/\/image.blockchain.news\/features\/45B7A801D37E36DC0019AE0310A0ED0160FBF51AC1E55381847ABA1D7FFAC0B0.jpg\" alt=\"Alibaba Unveils Qwen3.8-Flash-Next AI Model with 176B Parameters\" loading=\"eager\" width=\"1200\" height=\"630\"> <\/a> <\/figure>\n<p>Alibaba has released the weights for its Qwen3.8-Flash-Next multimodal model, a 176-billion parameter Mixture-of-Experts (MoE) architecture designed to handle long-context tasks such as agentic coding, document processing, and high-volume workflows. This release, announced on August 26, serves as a technical preview of the upcoming Qwen4 model family and emphasizes cost-efficiency and scalability for developers.<\/p>\n<h2>Technical Highlights<\/h2>\n<p>Qwen3.8-Flash-Next combines two key innovations for long-context inference: Gated DeltaNet (GDN) and Qwen Sparse Attention (QSA). GDN compresses historical context into a fixed-size recurrent state, avoiding the memory growth typically associated with long sequences. Meanwhile, QSA works at the micro-block level to reduce computational overhead while maintaining retrieval precision. Together, these mechanisms enable efficient processing of sequences up to 1 million tokens, an area where traditional attention models struggle.<\/p>\n<p>Benchmarks shared by Alibaba indicate QSA delivers tangible performance gains. For example, in a 1-million-token workload with a 90% prefix-cache hit rate, Qwen3.8 achieved 8.6x the prefill throughput of its predecessor, Qwen3.7-Plus. During prefill and decoding phases, the sparse attention kernel outperformed full attention by 7.6x and 4.9x, respectively, underscoring its utility for high-load applications.<\/p>\n<h2>Optimized for NVIDIA GB300 NVL72<\/h2>\n<p>NVIDIA has collaborated with Alibaba to validate Qwen3.8-Flash-Next on its GB300 NVL72 platform, which integrates 72 Blackwell Ultra GPUs with a 130 TB\/s NVLink communication fabric. The model achieves over 16,000 tokens per second per GPU and supports more than 200 tokens per second per user, making it suitable for large-scale deployment in production environments. Developers can also experiment with the model on smaller systems like NVIDIA DGX Spark clusters or workstations equipped with RTX PRO 6000 Blackwell GPUs.<\/p>\n<h2>Developer Tools and Accessibility<\/h2>\n<p>Qwen3.8-Flash-Next is available as open weights, and developers can download it from platforms like Hugging Face and ModelScope. NVIDIA supports fine-tuning through its NeMo AutoModel library, enabling domain-specific adaptations via efficient methods such as LoRA. Reinforcement learning workflows are also supported using NeMo RL recipes.<\/p>\n<p>For inference, developers can choose from multiple tools, including SGLang, vLLM, and NVIDIA TensorRT, ensuring flexibility in deployment. Alibaba has positioned this release as a cost-efficient model for developers to prototype and scale workflows ahead of the Qwen4 launch.<\/p>\n<h2>Why This Matters<\/h2>\n<p>Qwen3.8-Flash-Next represents a significant milestone in the evolution of large language models. By addressing the scalability challenges of long-context tasks, Alibaba is laying the groundwork for broader adoption of multimodal AI solutions. While this release is primarily a technical preview, its innovations in sparse attention and memory efficiency could influence the next generation of AI architectures and tools.<\/p>\n<p><span><i>Image source: Shutterstock<\/i><\/span> <!-- Divider --> <!-- Bookmark button --> <!-- Bookmark button END --> <!-- Author info END --> <!-- Divider --> <a href=\"https:\/\/blockchain.news\/\">Source<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Darius Baruo Aug 26, 2026 17:39 Alibaba releases Qwen3.8-Flash-Next, a 176B-parameter Mixture-of-Experts model, previewing Qwen4 architecture and optimized for long-context tasks. Alibaba has released the weights for its Qwen3.8-Flash-Next multimodal model, a 176-billion parameter Mixture-of-Experts (MoE) architecture designed to handle long-context tasks such as agentic coding, document processing, and high-volume workflows. This release, announced on [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":648844,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12],"tags":[1129,2330,23562,25,2148,26428],"class_list":{"0":"post-648843","1":"post","2":"type-post","3":"status-publish","4":"format-standard","5":"has-post-thumbnail","7":"category-blockchain","8":"tag-ai","9":"tag-alibaba","10":"tag-mixture-of-experts","11":"tag-news","12":"tag-nvidia","13":"tag-qwen3-8"},"_links":{"self":[{"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/posts\/648843","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/comments?post=648843"}],"version-history":[{"count":0,"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/posts\/648843\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/media\/648844"}],"wp:attachment":[{"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/media?parent=648843"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/categories?post=648843"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/tags?post=648843"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}