<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>AMD on Matt Suiche</title><link>https://www.msuiche.com/tags/amd/</link><description>Recent content in AMD on Matt Suiche</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Wed, 15 Oct 2025 02:00:00 +0000</lastBuildDate><atom:link href="https://www.msuiche.com/tags/amd/index.xml" rel="self" type="application/rss+xml"/><item><title>AMD GPU Support in Triton Gluon Framework</title><link>https://www.msuiche.com/posts/amd-gpu-support-in-triton-gluon-framework/</link><pubDate>Wed, 15 Oct 2025 02:00:00 +0000</pubDate><guid>https://www.msuiche.com/posts/amd-gpu-support-in-triton-gluon-framework/</guid><description>&lt;h2 id="introduction"&gt;Introduction&lt;a href="#introduction" class="anchor" aria-label="Link to Introduction"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;This document analyzes AMD GPU support implementation in Triton&amp;rsquo;s Gluon framework, examining architecture-specific optimizations, performance characteristics, and implementation details relative to NVIDIA GPU support.&lt;/p&gt;
&lt;p&gt;For background on Gluon and its motivation as a lower-level alternative to Triton, see my previous post: &lt;a href="https://www.msuiche.com/posts/gluon-when-triton-isnt-low-level-enough/" target="_blank" rel="noopener"&gt;&amp;ldquo;Gluon: When Triton Isn&amp;rsquo;t Low-Level Enough&amp;rdquo;&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="background-gpu-programming-architecture-landscape"&gt;Background: GPU Programming Architecture Landscape&lt;a href="#background-gpu-programming-architecture-landscape" class="anchor" aria-label="Link to Background: GPU Programming Architecture Landscape"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;The GPU programming ecosystem has evolved with distinct architectural approaches between NVIDIA and AMD, creating implementation challenges for cross-platform frameworks.&lt;/p&gt;</description></item><item><title>Multi-GPU Programming with AMD's Iris Framework for Triton</title><link>https://www.msuiche.com/posts/multi-gpu-programming-with-amds-iris-framework-for-triton/</link><pubDate>Sun, 28 Sep 2025 00:00:00 +0000</pubDate><guid>https://www.msuiche.com/posts/multi-gpu-programming-with-amds-iris-framework-for-triton/</guid><description>&lt;p&gt;GPU production constraints are creating infrastructure bottlenecks. Multi-GPU programming, particularly vendor-agnostic implementations, has become essential. In their &lt;a href="https://www.youtube.com/watch?v=H2bzSn5ZPks" target="_blank" rel="noopener"&gt;GPU Mode presentation&lt;/a&gt;, AMD Research engineers Muhammad Awad, Muhammad Osama, and Brandon Potter introduced Iris—a Python library that enables fine-grained multi-GPU programming in Triton. Similarly to my previous &lt;a href="https://www.msuiche.com/posts/gluon-when-triton-isnt-low-level-enough/"&gt;Gluon blogpost&lt;/a&gt;, this post captures my understanding and interpretation of their work, serving as both technical documentation and personal reference for this emerging multi-GPU programming paradigm.&lt;/p&gt;
&lt;h2 id="technical-problem"&gt;Technical Problem&lt;a href="#technical-problem" class="anchor" aria-label="Link to Technical Problem"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Current multi-GPU programming uses bulk synchronous models (BSP) through libraries like NCCL. This model enforces sequential phases:&lt;/p&gt;</description></item></channel></rss>