{"id":463997,"date":"2024-09-28T07:13:27","date_gmt":"2024-09-28T07:13:27","guid":{"rendered":"https:\/\/Blockchain.News\/news\/amd-introduces-amd-135m-breakthrough-small-language-models"},"modified":"2024-09-28T07:13:27","modified_gmt":"2024-09-28T07:13:27","slug":"amd-introduces-amd-135m-a-breakthrough-in-small-language-models","status":"publish","type":"post","link":"https:\/\/e-bitco.in\/index.php\/2024\/09\/28\/amd-introduces-amd-135m-a-breakthrough-in-small-language-models\/","title":{"rendered":"AMD Introduces AMD-135M: A Breakthrough in Small Language Models"},"content":{"rendered":"<figure class=\"figure mt-2\">\n<p> <a href=\"https:\/\/blockchain.news\/Profile\/Luisa-Crawford\">Luisa Crawford<\/a> <span class=\"publication-date ml-2\"> Sep 28, 2024 07:13<\/span> <\/p>\n<p class=\"lead\">AMD has unveiled its first small language model, AMD-135M, with Speculative Decoding, enhancing AI model efficiency and performance.<\/p>\n<p> <a href=\"https:\/\/blockchainstock.azureedge.net:443\/features\/4535012436EAA292C1BF9F56896B89FB79C06A2CF3E57F22F0C60AD08C28AD78.jpg\"> <img decoding=\"async\" class=\"rounded\" src=\"https:\/\/blockchainstock.azureedge.net:443\/features\/4535012436EAA292C1BF9F56896B89FB79C06A2CF3E57F22F0C60AD08C28AD78.jpg\" alt=\"AMD Introduces AMD-135M: A Breakthrough in Small Language Models\"> <\/a> <\/figure>\n<p>In a significant development within the artificial intelligence sector, AMD has announced the release of its first small language model (SLM), AMD-135M. This new model aims to offer specialized capabilities while addressing some of the limitations faced by large language models (LLMs) such as GPT-4 and Llama, according to <a rel=\"nofollow\" href=\"https:\/\/community.amd.com\/t5\/ai\/amd-unveils-its-first-small-language-model-amd-135m\/ba-p\/711368\">AMD.com<\/a>.<\/p>\n<h2>AMD-135M: First AMD Small Language Model<\/h2>\n<p>The AMD-135M, part of the Llama family, is AMD&#8217;s pioneering effort in the SLM arena. The model was trained from scratch using AMD Instinct\u2122 MI250 accelerators and 670 billion tokens. The training process resulted in two distinct models: AMD-Llama-135M and AMD-Llama-135M-code. The former underwent pretraining with general data, while the latter was fine-tuned with an additional 20 billion tokens specifically for code data.<\/p>\n<p><strong>Pretraining<\/strong>: AMD-Llama-135M was trained over six days using four MI250 nodes. The code-focused variant, AMD-Llama-135M-code, required an additional four days for fine-tuning.<\/p>\n<p>All associated training code, datasets, and model weights are open-sourced, enabling developers to reproduce the model and contribute to the training of other SLMs and LLMs.<\/p>\n<h2>Optimization with Speculative Decoding<\/h2>\n<p>One of the notable advancements in AMD-135M is the use of speculative decoding. Traditional autoregressive approaches in large language models often suffer from low memory access efficiency, as each forward pass generates only a single token. Speculative decoding addresses this by employing a small draft model to generate candidate tokens, which are then verified by a larger target model. This method allows multiple tokens to be generated per forward pass, significantly improving memory access efficiency and inference speed.<\/p>\n<h2>Inference Performance Acceleration<\/h2>\n<p>AMD has tested the performance of AMD-Llama-135M-code as a draft model for CodeLlama-7b on various hardware configurations, including the MI250 accelerator and the Ryzen\u2122 AI processor. The results indicated a considerable speedup in inference performance when speculative decoding was employed. This enhancement establishes an end-to-end workflow for training and inferencing on selected AMD platforms.<\/p>\n<h2>Next Steps<\/h2>\n<p>By providing an open-source reference implementation, AMD aims to foster innovation within the AI community. The company encourages developers to explore and contribute to this new frontier in AI technology.<\/p>\n<p>For more details on AMD-135M, visit the full technical blog on <a rel=\"nofollow\" href=\"https:\/\/community.amd.com\/t5\/ai\/amd-unveils-its-first-small-language-model-amd-135m\/ba-p\/711368\">AMD.com<\/a>.<\/p>\n<p><span><i>Image source: Shutterstock<\/i><\/span> <!-- Divider --> <!-- Author info END --> <!-- Divider --> <a href=\"https:\/\/blockchain.news\/\">Source<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Luisa Crawford Sep 28, 2024 07:13 AMD has unveiled its first small language model, AMD-135M, with Speculative Decoding, enhancing AI model efficiency and performance. In a significant development within the artificial intelligence sector, AMD has announced the release of its first small language model (SLM), AMD-135M. This new model aims to offer specialized capabilities while [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":463998,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12],"tags":[1129,13902,16729,25],"class_list":{"0":"post-463997","1":"post","2":"type-post","3":"status-publish","4":"format-standard","5":"has-post-thumbnail","7":"category-blockchain","8":"tag-ai","9":"tag-amd","10":"tag-language-models","11":"tag-news"},"_links":{"self":[{"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/posts\/463997","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/comments?post=463997"}],"version-history":[{"count":0,"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/posts\/463997\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/media\/463998"}],"wp:attachment":[{"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/media?parent=463997"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/categories?post=463997"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/e-bitco.in\/index.php\/wp-json\/wp\/v2\/tags?post=463997"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}