{"id":502892,"date":"2025-10-20T15:21:43","date_gmt":"2025-10-20T15:21:43","guid":{"rendered":"https:\/\/Blockchain.News\/news\/nvidia-nvl72-revolutionizing-moe-model-scaling"},"modified":"2025-10-20T15:21:43","modified_gmt":"2025-10-20T15:21:43","slug":"nvidia-nvl72-revolutionizing-moe-model-scaling-with-expert-parallelism","status":"publish","type":"post","link":"https:\/\/e-bitco.in\/index.php\/2025\/10\/20\/nvidia-nvl72-revolutionizing-moe-model-scaling-with-expert-parallelism\/","title":{"rendered":"NVIDIA NVL72: Revolutionizing MoE Model Scaling with Expert Parallelism"},"content":{"rendered":"<figure class=\"figure mt-2\">\n<p> <a href=\"https:\/\/blockchain.news\/Profile\/Joerg-Hiller\">Joerg Hiller<\/a> <span class=\"publication-date ml-2\"> Oct 20, 2025 15:21<\/span> <\/p>\n<p class=\"lead\">NVIDIA&#8217;s NVL72 systems are transforming large-scale MoE model deployment by introducing Wide Expert Parallelism, optimizing performance and reducing costs.<\/p>\n<p> <a href=\"https:\/\/image.blockchain.news:443\/features\/D8E08E86F8EDBDDCD68414CF49BDD8B1401B11A69515DFF98E6B2B03EE9CF9D7.jpg\"> <img decoding=\"async\" class=\"rounded\" src=\"https:\/\/image.blockchain.news:443\/features\/D8E08E86F8EDBDDCD68414CF49BDD8B1401B11A69515DFF98E6B2B03EE9CF9D7.jpg\" alt=\"NVIDIA NVL72: Revolutionizing MoE Model Scaling with Expert Parallelism\"> <\/a> <\/figure>\n<p>NVIDIA is advancing the deployment of large-scale Mixture of Experts (MoE) models with its NVL72 rack-scale systems, leveraging Wide Expert Parallelism (Wide-EP) to optimize performance and reduce costs, according to <a rel=\"nofollow\" href=\"https:\/\/developer.nvidia.com\/blog\/scaling-large-moe-models-with-wide-expert-parallelism-on-nvl72-rack-scale-systems\/\">NVIDIA&#8217;s blog<\/a>. This approach addresses the challenges of scaling MoE architectures, which are more efficient than dense models by activating only a subset of trained parameters per token.<\/p>\n<h2>Expert Parallelism and Its Impact<\/h2>\n<p>Expert Parallelism (EP) strategically distributes MoE model experts across multiple GPUs, enhancing computation and memory bandwidth utilization. As models like DeepSeek-R1 expand to hundreds of billions of parameters, EP becomes crucial for maintaining high performance and reducing memory pressure.<\/p>\n<p>Large-scale EP, which distributes experts across numerous GPUs, increases bandwidth and supports larger batch sizes, improving GPU utilization. However, it introduces new system-level constraints, which NVIDIA&#8217;s TensorRT-LLM Wide-EP aims to address through algorithmic optimizations targeting compute and memory bottlenecks.<\/p>\n<h2>System Design and Architecture<\/h2>\n<p>The effectiveness of scaling EP relies heavily on system design and architecture, particularly the interconnect bandwidth and topology, which facilitate efficient memory movement and communication. NVIDIA&#8217;s NVL72 systems use optimized software and kernels to manage expert-to-expert traffic, ensuring practical and efficient large-scale EP deployment.<\/p>\n<h3>Addressing Communication Overhead<\/h3>\n<p>Communication overhead is a significant challenge in large-scale EP, particularly during the inference decode phase when distributed experts must exchange information. NVIDIA&#8217;s NVLink technology, with its 130 TB\/s aggregate bandwidth, plays a crucial role in mitigating these overheads, making large-scale EP feasible.<\/p>\n<h3>Kernel Optimization and Load Balancing<\/h3>\n<p>To optimize expert routing, custom communication kernels are implemented to manage non-static data sizes effectively. NVIDIA&#8217;s Expert Parallel Load Balancer (EPLB) further enhances load balancing by redistributing experts to prevent over- or under-utilization of GPUs, crucial for maintaining efficiency in real-time production systems.<\/p>\n<h2>Implications for AI Inference<\/h2>\n<p>Wide-EP on NVIDIA&#8217;s NVL72 systems provides a scalable solution for MoE models, reducing weight-loading pressure and improving GroupGEMM efficiency. In testing, large EP configurations demonstrated up to 1.8x higher per-GPU throughput compared to smaller setups, highlighting the potential for significant performance gains.<\/p>\n<p>The advancements in Wide-EP not only improve throughput and latency but also enhance system economics by increasing concurrency and GPU efficiency. This positions NVIDIA&#8217;s NVL72 as a pivotal player in the cost-effective deployment of trillion-parameter models, offering developers, researchers, and infrastructure teams new opportunities to optimize AI workloads.<\/p>\n<p><span><i>Image source: Shutterstock<\/i><\/span> <!-- Divider --> <!-- Author info END --> <!-- Divider --> <a href=\"https:\/\/blockchain.news\/\">Source<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Joerg Hiller Oct 20, 2025 15:21 NVIDIA&#8217;s NVL72 systems are transforming large-scale MoE model deployment by introducing Wide Expert Parallelism, optimizing performance and reducing costs. NVIDIA is advancing the deployment of large-scale Mixture of Experts (MoE) models with its NVL72 rack-scale systems, leveraging Wide Expert Parallelism (Wide-EP) to optimize performance and reduce costs, according to [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":502893,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12],"tags":[1129,22693,22692,25,2148],"class_list":{"0":"post-502892","1":"post","2":"type-post","3":"status-publish","4":"format-standard","5":"has-post-thumbnail","7":"category-blockchain","8":"tag-ai","9":"tag-expert-parallelism","10":"tag-moe-models","11":"tag-news","12":"tag-nvidia"},"_links":{"self":[{"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/posts\/502892","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/comments?post=502892"}],"version-history":[{"count":0,"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/posts\/502892\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/media\/502893"}],"wp:attachment":[{"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/media?parent=502892"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/categories?post=502892"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/tags?post=502892"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}