AMD GPUs (ROCm)

GeoSWE runs on AMD GPUs through CuPy’s ROCm build. The API, the kernels and the SWE_* settings are the same as on NVIDIA; geoswe.gpu_platform() returns "hip" instead of "cuda".

What has been tested

AMD Instinct MI210 and MI250X (gfx90a) on OLCF Frontier, with ROCm 7.0.2 and 7.2.0 and two CuPy builds: cupy-rocm-7-0 14.2.0 from PyPI and AMD’s amd-cupy 13.5.1. Both pass the whole GPU test suite. Other AMD GPUs and ROCm versions are untested. CuPy itself still labels its ROCm support experimental.

Installing

CuPy’s ROCm build compiles kernels with the ROCm installation on the machine, so ROCm 7 must be installed and visible at run time:

  • hipcc on PATH (CuPy runs it to find the include directories), and

  • ROCM_HOME pointing at the ROCm tree, for example /opt/rocm-7.2.0.

pip install "geoswe[gpu-rocm]"                       # CuPy for ROCm 7 from PyPI
pip install "geoswe[gpu-rocm,mpi,io,forcings]"       # with MPI, GeoTIFF I/O and forcings

AMD also publishes its own build, which OLCF documents for Frontier. Install it instead of the extra, never next to it:

pip install amd-cupy --extra-index-url https://pypi.amd.com/rocm-7.2.0/simple
pip install "geoswe[mpi,io,forcings]"

Check the install:

python -c "import geoswe; print(geoswe.get_backend(), geoswe.gpu_platform())"   # cupy hip
pytest -m gpu -q

The first run compiles every kernel, which takes a minute or so; later runs load them from CuPy’s cache. On a cluster, put that cache on a filesystem the compute nodes share, and keep AMD’s own compiler cache off NFS inside jobs:

export CUPY_CACHE_DIR=/path/on/scratch/cupy_cache       # default: ~/.cupy/kernel_cache
export AMD_COMGR_CACHE_DIR=/tmp/$USER-comgr             # default: ~/.cache/comgr

What differs from NVIDIA

The kernels are CUDA C. NVIDIA’s compiler (NVRTC) and AMD’s (clang, through HIP) disagree about three things in them, and geoswe.backend handles each one. On NVIDIA nothing changes: kernel sources and options reach CuPy exactly as before.

NVIDIA

AMD

What GeoSWE does on AMD

Register cap -maxrregcount (SWE_DENSE_MAXRREG, SWE_FLAT_MAXRREG)

applied on H100

not an option of the compiler

never requested; ignored with a warning if set by hand

__fmul_rn / __fadd_rn

never fused into a multiply-add

plain * and +, which clang fuses

redefined with contraction switched off

Warp-level compaction

32 lanes, 32-bit masks

64 lanes, 64-bit masks

a kernel without warp intrinsics

The first one also needed a fix to device detection: CuPy reports the compute capability of a gfx90a card as "90", the same string as an H100.

Floating-point contraction

GEOSWE_HIP_FP_CONTRACT sets how the ROCm compiler may fuse a*b + c into one multiply-add:

Value

Meaning

off (default)

never; every operation rounds once, in source order

on

only inside a single source expression

fast

at the optimizer’s discretion, across statements (clang’s own default for HIP)

GeoSWE relies on several code paths giving identical bits: the fused and the split time step, the dense and the compressed mesh, a run and its restart. With fast the same expression can be fused in one kernel and not in another, and the fused and split steps with sub-grid storage then differ in the last bit (tests/test_gpu_storage_fused.py fails). With off or on the arithmetic is fixed by the source text.

off is the default because it makes those identities hold by construction: two kernels agree whenever their statements do. It is the only mode in which the dense and the compressed solver agree bit for bit on the bowl-and-dam problem of tests/test_gpu_compressed_equiv.py; on and fast leave one unit in the last place between them. on is about 3 % faster (25.0 against 25.8 ms/step on an MI250X at 147 M cells; fast is in between) and keeps the fused and split steps identical in the test suite. All three modes keep a lake at rest to round-off.

Results on AMD and NVIDIA are not bit-identical to each other: the compilers and their math libraries differ, at round-off level.

Multiple GPUs

One MPI rank drives one GPU, as on NVIDIA. Under Slurm, let the scheduler bind the devices:

srun -n 8 --gpus-per-task=1 --gpu-bind=closest python my_run.py

The halo travels through host memory unless SWE_HALO_CUDA_AWARE=1 asks for GPU-aware MPI. Despite the name, that setting is not specific to CUDA: CuPy’s ROCm build exposes device arrays through the same interface, and mpi4py hands the pointers to MPI. With HPE Cray MPICH this needs two more things:

  • mpi4py built with the craype-accel-amd-* module loaded, so that it links the GPU transport library, and

  • MPICH_GPU_SUPPORT_ENABLED=1 in the job.

If SWE_HALO_CUDA_AWARE=1 is set while Cray MPICH runs without its GPU support, GeoSWE warns and stages through the host instead of passing device pointers to a library that would treat them as host memory.

OLCF Frontier

benchmark/frontier_amd/ in the repository has the module set, the build steps for a GPU-aware mpi4py, and a one-node job that runs the GPU test suite, a partition-invariance check and a weak-scaling sweep. Measured there on one node (an MI250X card is two devices, so eight per node; 147 M cells per device, float32, host-staged halo):

Configuration

1 device

8 devices

Weak efficiency

dense

25.8 ms/step

26.1 ms/step

99.0 %

flat-full

23.7 ms/step

23.9 ms/step

99.2 %

Repeat launches differ by up to 2 %. The global solution is bit-identical on 1, 2, 4 and 8 devices, with the host-staged and the GPU-aware halo, with and without the halo/compute overlap.

On two nodes, the scaling harness of benchmark/scaling_640m (640 M cells per device, in the configuration its launchers pin) gives:

Configuration

Weak scaling, 16 devices, 10.24 B cells

Strong scaling, 16 devices

flat-full

118.7 ms/step, 99.2 % efficiency

15.0x

dense

190.2 ms/step, 99.6 % efficiency

15.4x

Both benchmarks keep every cell active, so they are the flat-full and dense configurations of the three; the terrain-selected flat-active configuration is the production one. With the solver’s defaults the same 10.24 billion cells take 102 ms/step (flat-full) and 112 ms/step (dense). The solution on sixteen devices across the two nodes has the digest of the one-device run.

Note

A device here is one GCD, half an MI250X card. For the same 640 M cells with the solver’s defaults, the H100 of the scaling figure in the repository README takes about 42 ms/step (flat) and 46 ms/step (dense), where one GCD takes 101 and 111. That is 2.4 times faster, the ratio of their FP32 lanes (16,896 CUDA cores to 7,040 stream processors), so a whole MI250X card delivers about 0.83 of an H100. The machines differ in more than the GPU: read this as arithmetic on published numbers, not as a controlled comparison.

Not done yet

  • Of the paper’s benchmark cases, only the synthetic scaling harness has been run on AMD hardware.

  • The register cap that speeds up the residual kernel on H100 has no ROCm counterpart, and no AMD-specific tuning of block sizes has been tried.

  • The scaling_640m harness has been run on up to two nodes (16 devices); the weak-scaling campaign at one billion cells per GCD (run_weak_1b.sbatch) reaches 128 nodes, 1024 devices and 1.024 trillion cells.