AMD GPUs (ROCm)¶
GeoSWE runs on AMD GPUs through CuPy’s ROCm build. The API, the kernels and the
SWE_* settings are the same as on NVIDIA; geoswe.gpu_platform() returns
"hip" instead of "cuda".
What has been tested
AMD Instinct MI210 and MI250X (gfx90a) on OLCF Frontier, with ROCm 7.0.2 and
7.2.0 and two CuPy builds: cupy-rocm-7-0 14.2.0 from PyPI and AMD’s amd-cupy
13.5.1. Both pass the whole GPU test suite. Other AMD GPUs and ROCm versions are
untested. CuPy itself still labels its ROCm support experimental.
Installing¶
CuPy’s ROCm build compiles kernels with the ROCm installation on the machine, so ROCm 7 must be installed and visible at run time:
hipcconPATH(CuPy runs it to find the include directories), andROCM_HOMEpointing at the ROCm tree, for example/opt/rocm-7.2.0.
pip install "geoswe[gpu-rocm]" # CuPy for ROCm 7 from PyPI
pip install "geoswe[gpu-rocm,mpi,io,forcings]" # with MPI, GeoTIFF I/O and forcings
AMD also publishes its own build, which OLCF documents for Frontier. Install it instead of the extra, never next to it:
pip install amd-cupy --extra-index-url https://pypi.amd.com/rocm-7.2.0/simple
pip install "geoswe[mpi,io,forcings]"
Check the install:
python -c "import geoswe; print(geoswe.get_backend(), geoswe.gpu_platform())" # cupy hip
pytest -m gpu -q
The first run compiles every kernel, which takes a minute or so; later runs load them from CuPy’s cache. On a cluster, put that cache on a filesystem the compute nodes share, and keep AMD’s own compiler cache off NFS inside jobs:
export CUPY_CACHE_DIR=/path/on/scratch/cupy_cache # default: ~/.cupy/kernel_cache
export AMD_COMGR_CACHE_DIR=/tmp/$USER-comgr # default: ~/.cache/comgr
What differs from NVIDIA¶
The kernels are CUDA C. NVIDIA’s compiler (NVRTC) and AMD’s (clang, through HIP)
disagree about three things in them, and geoswe.backend handles each one. On
NVIDIA nothing changes: kernel sources and options reach CuPy exactly as before.
NVIDIA |
AMD |
What GeoSWE does on AMD |
|
|---|---|---|---|
Register cap |
applied on H100 |
not an option of the compiler |
never requested; ignored with a warning if set by hand |
|
never fused into a multiply-add |
plain |
redefined with contraction switched off |
Warp-level compaction |
32 lanes, 32-bit masks |
64 lanes, 64-bit masks |
a kernel without warp intrinsics |
The first one also needed a fix to device detection: CuPy reports the compute
capability of a gfx90a card as "90", the same string as an H100.
Floating-point contraction¶
GEOSWE_HIP_FP_CONTRACT sets how the ROCm compiler may fuse a*b + c into one
multiply-add:
Value |
Meaning |
|---|---|
|
never; every operation rounds once, in source order |
|
only inside a single source expression |
|
at the optimizer’s discretion, across statements (clang’s own default for HIP) |
GeoSWE relies on several code paths giving identical bits: the fused and the
split time step, the dense and the compressed mesh, a run and its restart. With
fast the same expression can be fused in one kernel and not in another, and the
fused and split steps with sub-grid storage then differ in the last bit
(tests/test_gpu_storage_fused.py fails). With off or on the arithmetic is
fixed by the source text.
off is the default because it makes those identities hold by construction: two
kernels agree whenever their statements do. It is the only mode in which the
dense and the compressed solver agree bit for bit on the bowl-and-dam problem of
tests/test_gpu_compressed_equiv.py; on and fast leave one unit in the last
place between them. on is about 3 % faster (25.0 against 25.8 ms/step on an
MI250X at 147 M cells; fast is in between) and keeps the fused and split steps
identical in the test suite. All three modes keep a lake at rest to round-off.
Results on AMD and NVIDIA are not bit-identical to each other: the compilers and their math libraries differ, at round-off level.
Multiple GPUs¶
One MPI rank drives one GPU, as on NVIDIA. Under Slurm, let the scheduler bind the devices:
srun -n 8 --gpus-per-task=1 --gpu-bind=closest python my_run.py
The halo travels through host memory unless SWE_HALO_CUDA_AWARE=1 asks for
GPU-aware MPI. Despite the name, that setting is not specific to CUDA: CuPy’s
ROCm build exposes device arrays through the same interface, and mpi4py hands
the pointers to MPI. With HPE Cray MPICH this needs two more things:
mpi4pybuilt with thecraype-accel-amd-*module loaded, so that it links the GPU transport library, andMPICH_GPU_SUPPORT_ENABLED=1in the job.
If SWE_HALO_CUDA_AWARE=1 is set while Cray MPICH runs without its GPU support,
GeoSWE warns and stages through the host instead of passing device pointers to a
library that would treat them as host memory.
OLCF Frontier¶
benchmark/frontier_amd/ in the repository has the module set, the build steps
for a GPU-aware mpi4py, and a one-node job that runs the GPU test suite, a
partition-invariance check and a weak-scaling sweep. Measured there on one node
(an MI250X card is two devices, so eight per node; 147 M cells per device,
float32, host-staged halo):
Configuration |
1 device |
8 devices |
Weak efficiency |
|---|---|---|---|
dense |
25.8 ms/step |
26.1 ms/step |
99.0 % |
flat-full |
23.7 ms/step |
23.9 ms/step |
99.2 % |
Repeat launches differ by up to 2 %. The global solution is bit-identical on 1, 2, 4 and 8 devices, with the host-staged and the GPU-aware halo, with and without the halo/compute overlap.
On two nodes, the scaling harness of benchmark/scaling_640m (640 M cells per
device, in the configuration its launchers pin) gives:
Configuration |
Weak scaling, 16 devices, 10.24 B cells |
Strong scaling, 16 devices |
|---|---|---|
flat-full |
118.7 ms/step, 99.2 % efficiency |
15.0x |
dense |
190.2 ms/step, 99.6 % efficiency |
15.4x |
Both benchmarks keep every cell active, so they are the flat-full and dense configurations of the three; the terrain-selected flat-active configuration is the production one. With the solver’s defaults the same 10.24 billion cells take 102 ms/step (flat-full) and 112 ms/step (dense). The solution on sixteen devices across the two nodes has the digest of the one-device run.
Note
A device here is one GCD, half an MI250X card. For the same 640 M cells with the solver’s defaults, the H100 of the scaling figure in the repository README takes about 42 ms/step (flat) and 46 ms/step (dense), where one GCD takes 101 and 111. That is 2.4 times faster, the ratio of their FP32 lanes (16,896 CUDA cores to 7,040 stream processors), so a whole MI250X card delivers about 0.83 of an H100. The machines differ in more than the GPU: read this as arithmetic on published numbers, not as a controlled comparison.
Not done yet¶
Of the paper’s benchmark cases, only the synthetic scaling harness has been run on AMD hardware.
The register cap that speeds up the residual kernel on H100 has no ROCm counterpart, and no AMD-specific tuning of block sizes has been tried.
The
scaling_640mharness has been run on up to two nodes (16 devices); the weak-scaling campaign at one billion cells per GCD (run_weak_1b.sbatch) reaches 128 nodes, 1024 devices and 1.024 trillion cells.