Skip to content

Repository files navigation

Iris: First-Class Multi-GPU Programming Experience in Triton

License: MITRuffIris TestsDOIDOI

Iris is a Triton-based framework for Remote Memory Access (RMA) operations developed by AMD's Research and Advanced Development team. Iris provides SHMEM-like APIs within Triton for Multi-GPU programming. Iris' goal is to make Multi-GPU programming a first-class citizen in Triton while retaining Triton's programmability and performance.

Latest with Iris 🔥

Key Features

  • SHMEM-like RMA: Iris provides SHMEM-like RMA support in Triton.
  • Simple and Intuitive API: Iris provides simple and intuitive RMA APIs. Writing multi-GPU programs is as easy as writing single-GPU programs.
  • Triton-based: Iris is built on top of Triton and inherits Triton's performance and capabilities.
  • Triton Gluon-based backend (Experimental): Includes an optional backend built on Triton’s Gluon language, a lower-level GPU programming model that exposes explicit control over layouts, memory, and data movement—ideal for users seeking maximal performance and hardware-level optimization.

Documentation

API Example

Here's a simple example showing how to perform remote memory operations between GPUs using Iris:

importtorchimporttorch.distributedasdistimporttorch.multiprocessingasmpimporttritonimporttriton.languageastlimportiris# Device-side APIs@triton.jitdefkernel(buffer, buffer_size: tl.constexpr, block_size: tl.constexpr, heap_bases_ptr):
# Compute start index of this blockpid=tl.program_id(0)
block_start=pid*block_sizeoffsets=block_start+tl.arange(0, block_size)
# Guard for out-of-bounds accessesmask=offsets<buffer_size# Store 1 in the target buffer at each offsetsource_rank=0target_rank=1iris.store(buffer+offsets, 1,
source_rank, target_rank,
heap_bases_ptr, mask=mask)
def_worker(rank, world_size):
# Torch distributed initializationdevice_id=rank%torch.cuda.device_count()
dist.init_process_group(
backend="nccl",
rank=rank,
world_size=world_size,
init_method="tcp://127.0.0.1:29500",
device_id=torch.device(f"cuda:{device_id}")
)
# Iris initializationheap_size=2**30# 1GiB symmetric heap for inter-GPU communicationiris_ctx=iris.iris(heap_size)
cur_rank=iris_ctx.get_rank()
# Iris tensor allocationbuffer_size=4096# 4K elements bufferbuffer=iris_ctx.zeros(buffer_size, device="cuda", dtype=torch.float32)
# Launch the kernel on rank 0block_size=1024grid=lambdameta: (triton.cdiv(buffer_size, meta["block_size"]),)
source_rank=0ifcur_rank==source_rank:
kernel[grid](
buffer,
buffer_size,
block_size,
iris_ctx.get_heap_bases(),
)
# Synchronize all ranksiris_ctx.barrier()
dist.destroy_process_group()
if__name__=="__main__":
world_size=2# Using two ranksmp.spawn(_worker, args=(world_size,), nprocs=world_size, join=True)

Gluon-style API (Experimental)

Iris also provides an experimental cleaner API using Triton's Gluon with @gluon.jit decorator:

Note

Requirements for Gluon backend: ROCm 7.0+ and Triton commit aafec417bded34db6308f5b3d6023daefae43905 or later are required to use the experimental Gluon APIs.

importtorchimporttorch.distributedasdistimporttorch.multiprocessingasmpfromtriton.experimentalimportgluonfromtriton.experimental.gluonimportlanguageasglimportirisfromiris.gluonimportIrisDeviceCtx# Device-side APIs - context encapsulates heap_bases@gluon.jitdefkernel(IrisDeviceCtx: gl.constexpr, context_tensor,
buffer, buffer_size: gl.constexpr, block_size: gl.constexpr):
# Initialize device context from tensorctx=IrisDeviceCtx.initialize(context_tensor)
pid=gl.program_id(0)
block_start=pid*block_sizelayout: gl.constexpr=gl.BlockedLayout([1], [64], [1], [0])
offsets=block_start+gl.arange(0, block_size, layout=layout)
mask=offsets<buffer_size# Store 1 in the target buffer - no need to pass heap_bases separately!target_rank=1ctx.store(buffer+offsets, 1, target_rank, mask=mask)
def_worker(rank, world_size):
# Torch distributed initializationdevice_id=rank%torch.cuda.device_count()
dist.init_process_group(
backend="nccl",
rank=rank,
world_size=world_size,
init_method="tcp://127.0.0.1:29500",
device_id=torch.device(f"cuda:{device_id}")
)
# Iris initializationheap_size=2**30# 1GiB symmetric heapiris_ctx=iris.iris(heap_size)
context_tensor=iris_ctx.get_device_context() # Get encoded contextcur_rank=iris_ctx.get_rank()
# Iris tensor allocationbuffer_size=4096# 4K elements bufferbuffer=iris_ctx.zeros(buffer_size, device="cuda", dtype=torch.float32)
# Launch the kernel on rank 0block_size=1024grid= (buffer_size+block_size-1) //block_sizesource_rank=0ifcur_rank==source_rank:
kernel[(grid,)](IrisDeviceCtx, context_tensor,
buffer, buffer_size, block_size, num_warps=1)
# Synchronize all ranksiris_ctx.barrier()
dist.destroy_process_group()
if__name__=="__main__":
world_size=2# Using two ranksmp.spawn(_worker, args=(world_size,), nprocs=world_size, join=True)

Quick Start Guide

Quick Installation

Note

Requirements: Python 3.10+, PyTorch 2.0+ (ROCm version), ROCm 6.3.1+ HIP runtime, Triton, and setuptools>=61

For a quick installation directly from the repository:

pip install git+https://github.com/ROCm/iris.git

Docker Compose (Recommended for Development)

The recommended way to get started is using Docker Compose, which provides a development environment with the Iris directory mounted inside the container. This allows you to make changes to the code outside the container and see them reflected inside.

# Start the development container
docker compose up --build -d
# or depending on your docker version
docker-compose up --build -d
# Attach to the running container
docker attach iris-dev
# Install Iris in development modecd iris && pip install -e .

For baremetal install, Docker or Apptainer setup, see Installation.

Next Steps

Check out our examples directory for ready-to-run scripts and usage patterns, including peer-to-peer communication and GEMM benchmarks.

Supported GPUs

Iris currently supports:

  • MI300X, MI350X & MI355X

Note

Iris may work on other AMD GPUs with ROCm compatibility.

Roadmap

We plan to extend Iris with the following features:

  • Extended GPU Support: Testing and optimization for other AMD GPUs.
  • RDMA Support: Multi-node support using Remote Direct Memory Access (RDMA) for distributed computing across multiple machines.
  • End-to-End Integration: Comprehensive examples covering various use cases and end-to-end patterns.

Contributing

We welcome contributions! Please see our Contributing Guide for details on how to set up your development environment and contribute to the project.

Support

Need help? We're here to support you! Here are a few ways to get in touch:

  1. Open an Issue: Found a bug or have a feature request? Open an issue on GitHub
  2. Contact the Team: If GitHub issues aren't working for you or you need to reach us directly, feel free to contact our development team

We welcome your feedback and contributions!

How to Cite

If you use Iris or reference it in your research, please cite our work:

@misc{Awad:2025:IFM,
author = {Muhammad Awad and Muhammad Osama and Brandon Potter},
title = {Iris: First-Class Multi-{GPU} Programming Experience in {Triton}},
year = {2025},
archivePrefix = {arXiv},
eprint = {2511.12500},
primaryClass = {cs.DC},
doi = {10.48550/arXiv.2511.12500}
}
@misc{Trifan:2025:EMT,
author = {Octavian Alexandru Trifan and Karthik Sangaiah and Muhammad Awad and Muhammad Osama and Sumanth Gudaparthi and Alexandru Nicolau and Alexander Veidenbaum and Ganesh Dasika},
title = {Eliminating Multi-{GPU} Performance Taxes: A Systems Approach to Efficient Distributed {LLMs}},
year = {2025},
archivePrefix = {arXiv},
eprint = {2511.02168},
primaryClass = {cs.DC},
doi = {10.48550/arXiv.2511.02168}
}
@software{Awad:2025:IFM:Software,
author = {Muhammad Awad and Muhammad Osama and Brandon Potter},
title = {Iris: First-Class Multi-{GPU} Programming Experience in {Triton}},
year = 2025,
month = oct,
doi = {10.5281/zenodo.17382307},
url = {https://github.com/ROCm/iris}
}

License

This project is licensed under the MIT License - see the LICENSE file for details.

Releases

Used by

Contributors

Languages