Benchmarks and applications¶
GeoSWE has been benchmarked head-to-head against three established GPU shallow-water codes on audited-identical inputs, and exercised from county to continental scale.
Multi-GPU scaling¶
On a synthetic uniform domain (so the result isolates the solver and halo from
real-terrain load imbalance), both storage paths scale near-ideally. The sweep
was run on two configurations: 1–16 NVIDIA H100 GPUs across two eight-GPU
nodes, and 1–32 MIG 2g.48gb slices of RTX PRO 6000 Blackwell GPUs across two
sixteen-slice nodes, which reaches twice the cell count.
Weak scaling (fixed work per GPU, 640 M cells each): 99.5% efficiency for the compressed path at 16 H100 GPUs, reaching 10.24 billion cells, and 98.8% at 32 Blackwell MIG slices across two nodes, reaching 20.48 billion cells. Per-step time stays nearly flat with rank count, including across the node boundary, because the halo exchange plus the scalar
dtall-reduce is an \(O(1)\) overhead.Strong scaling (fixed total problem): 15.5× speedup at 16 H100 GPUs (96.8% efficiency) and 26.5× at 32 Blackwell slices across two nodes (82.9%), tapering once each rank holds too few cells (the surface-to-volume trade-off).
The flat-full configuration is faster per step than the dense one at constant per-rank work, while producing bit-identical residuals on the same cells. The margin depends on the machine, so take the number with its machine: 9 % on the hardware of this study, one H100 MIG slice at 640 M cells per rank (115.1 against 127.7 ms/step). The same two harnesses on one full L40S at 320 M cells per rank give 22.6 against 28.4 ms/step, a 26 % margin, and under the first campaign’s split-forcings flags (
SWE_FUSE_FORCINGS=0,CFL_RESAMPLE_EVERY=5) 37.7 against 66.0 ms/step, a factor 1.75. Both tiers are bandwidth-bound at these sizes, and the dense tier is the one that moves: it is more sensitive to the card and to which forcings are fused. (These runs keep every cell active, so they measure the layout, not the active mask; see the three configurations.)
Do not difference the two hardware configurations against each other for a per-device ratio: they are separate machines with different interconnects and MIG partitioning, not a controlled per-device comparison.
Reproduce the shape of these curves with examples/ex05_scaling_bench.py. See
multi-GPU & MPI.
Accuracy: four-code comparison¶
On a 214-million-cell, 3 m Pinellas County benchmark (a Hurricane-Helene standing-tide case and a rain-driven ×10 variant), GeoSWE was compared against three established GPU shallow-water codes (ORNL’s TRITON, SynxFlow of the HiPIMS lineage, and SERGHEI) on audited-identical inputs:
GeoSWE’s flat-full configuration reproduces its dense solution to 0.05 cm RMSE (CSI 1.000) with identical step counts on the standing-tide benchmark, and the flat-active configuration agrees to 1.36 cm within its active set: the active-cell mesh is faithful.
GeoSWE’s flat-active configuration has the lowest wall time and peak GPU memory of the tested configurations at every GPU count: 3.5–3.9× faster per step, with 1.5–1.9× less per-rank memory, than the next-fastest comparison code; GeoSWE dense is 1.6–1.7× faster than TRITON on the same cells. The dense-to-flat-active speedup is not the storage layout alone: it also carries kernel fusion, and flat-full is the configuration that isolates the layout. The four codes’ scored standing-tide fields agree to pairwise CSI 0.989–0.994.
On the rain-driven (×10 rainfall) stress test, GeoSWE and SynxFlow complete the hour and agree closely (CSI 0.990 at 0.3 m, 1.5 cm RMSE); TRITON’s fp32 run and SERGHEI’s run at the benchmark floor do not reach a valid end field. On a steady rained slope with a high-accuracy reference solution, GeoSWE stays within 1 % of the reference film depth where TRITON and SERGHEI depart by tens of percent. GeoSWE is the only one of the four demonstrated at the continental scales below.
Applications¶
Case |
Resolution |
Active cells |
Hardware |
|---|---|---|---|
Pinellas County, FL (Helene) |
3 m |
125.6 M |
1 GPU |
Florida (Helene) |
10 m |
1.78 B |
4 GPUs |
CONUS (Helene) |
30 m |
8.88 B |
single node, 8 × H100 |
The continental cases use the compressed active-cell mesh with build-once caching and checkpoint/resume, so a 72-hour simulation survives job-time limits by resuming from the last checkpoint. Florida’s forcing spans most of its domain; on CONUS, Helene’s heavy rain covers only the Southeast, so that run is a capacity and workflow envelope rather than a continental flood simulation. The Florida and CONUS pipelines are not included in this release; the Pinellas benchmark exercises the same solver paths.
Note
These results and the full methodology are described in the GeoSWE paper; see citing. The cross-code benchmark establishes consistency between independently developed solvers on identical inputs, not validation against observations, and the application runs are computational demonstrations under stated, uncalibrated settings, not validated flood hindcasts.