Configuration reference¶
Numerical choices live in the Config dataclass. Deployment
and performance switches live in environment variables (GEOSWE_*, with older
SWE_* names); none of them has to be set.
Config fields¶
Scheme¶
Field |
Default |
Meaning |
|---|---|---|
|
|
the shallow-water equations, and the only released value: |
|
|
|
|
|
|
|
|
|
|
|
Courant number for |
|
|
gravity |
Tip
Config() with no arguments is the production flood configuration of the
GeoSWE paper: pde="baseline", flux="hllc", recon="first", well_balanced=True, wb_method="srm", time="euler", cfl=0.5. Higher-order reconstructions need
well_balanced=False to show their order; the well-balanced face states are
first-order by design.
Wetting / drying¶
Field |
Default |
Meaning |
|---|---|---|
|
|
physics wet/dry depth threshold (auto-raised to |
|
|
CFL-only floor. The default couples it to |
Well-balanced source¶
Field |
Default |
Meaning |
|---|---|---|
|
|
well-balanced bed treatment (exact lake-at-rest); face states are first-order |
|
|
|
Boundaries¶
Field |
Default |
Meaning |
|---|---|---|
|
|
|
|
|
1D Dirichlet states |
Friction¶
Field |
Default |
Meaning |
|---|---|---|
|
|
|
|
|
scalar Manning roughness; a map is given with |
|
|
the map on the padded grid |
|
|
boost \(n\) above this speed; |
|
|
quadratic-\(\alpha\) point-implicit root (the default, and what every published run used). Set |
Forcings¶
Field |
Default |
Meaning |
|---|---|---|
|
|
constant uniform rainfall rate, in m/s |
|
|
a |
|
|
a |
Storage / precision¶
Field |
Default |
Meaning |
|---|---|---|
|
follows the backend |
|
|
|
|
|
|
sub-grid channel storage ( |
|
|
the time step that sets that depth; |
The time step itself is not a Config field: Solver2D.run(t_end, dt_max=...)
caps it, and under rain run also bounds it by the CFL step of the film the rain
lays down (see numerical methods).
Environment variables¶
A script that builds its solvers directly needs only Config and, to choose
CPU or GPU, GEOSWE_BACKEND. The other variables below exist for the benchmark
and application run scripts, for performance work, and for debugging; the
defaults are the configuration every published run used, so none of them has
to be set. Where a variable has a second name, the GEOSWE_ one wins if both
are set and the other is an accepted alias (the research tree’s name, usually
SWE_-prefixed; GA, PROFILE_MEM and SIGMA_FREE_CFL carry no prefix at
all). Boolean flags take 1/0.
Simulation and I/O¶
Change what is simulated or written. The Read in column names every module whose
behaviour the variable changes, so a dense Solver2D run is affected only by the rows
that name solver.py or io_geotiff.py; the rest belong to the compressed tier
(compressed_solver.py, including the run_cached replay path) or to the coastal
driver (geoswe.runlib).
Variable |
|
Default |
Effect |
Read in |
|---|---|---|---|---|
|
unset |
array backend: “cupy” (GPU, NVIDIA or AMD) or “numpy” (CPU). Unset: the GPU when CuPy and a device are present, otherwise NumPy |
|
|
|
|
stop the compressed step loop when the CFL time step collapses below this (s); 0 (default) = no floor. It raises naming the step, the time and the dt instead of grinding on to the job’s wall clock. |
|
|
|
|
|
0 = linearized point-implicit friction root; unset = quadratic root (every published run). The driver folds it into the |
|
|
|
|
1 (default) = Green-Ampt infiltration on in the driver; 0 = off |
|
|
|
|
Green-Ampt cumulative-infiltration cap (m) |
|
|
|
|
“uniform” (default), “wtcap” (water-table storage cap) or “ssurgo” (per-map-unit parameters from GEOSWE_GA_SSURGO) |
|
|
|
`` |
soil .npz for GEOSWE_GA_MODE=ssurgo |
|
|
|
|
Green-Ampt moisture deficit override |
|
|
|
|
Green-Ampt water-table depth (m); default 1.5 |
|
|
|
print one progress line after this many steps with no other progress line, so a compressed run whose step has collapsed still reports; the regular lines are gated on simulated time. 0 = off |
|
|
|
|
`` |
two initial still-water stages, |
|
|
|
|
the largest prescribed ring stage (m) accepted before a run refuses to start: a datum mismatch (IGLD vs NAVD88) otherwise drives the tide tens of metres high. Raise it for a real extreme event (a tsunami study). Read by both entry points, the gauge CSVs the driver loads and a cached stage table |
|
|
|
`` |
1 = switch rainfall off |
|
|
|
`` |
gridded rainfall table (.npz) for the compressed replay path |
|
|
|
|
multiply the rainfall table by this factor (the x10 stress test); default 1 |
|
|
|
|
1 (default) = hold a two-row device window of the rainfall table; 0 = fully resident table. Bit-identical |
|
|
|
`` |
replace the rainfall table with a uniform rate (mm/h) |
|
|
|
how the compressed mesh closes its ghost ring. |
|
|
|
|
bed elevation dividing water from land for GEOSWE_RING_BC=hybrid (default: the ring stage) |
|
|
|
`` |
ambient still-water stage (m) imposed on the ghost ring; unset = ring stays dry |
|
|
|
|
0 (default) = write the compressed mesh’s depth frames as float16, whose step at a depth of 10 m is 0.0078 m; 1 = float32. The dense driver’s |
|
|
|
`` |
reserved (research-tree ring-stage ramp). Not implemented here: this solver imposes a constant ring stage, so any value raises |
|
|
|
`` |
1 = ring cells take the interior velocity (GEOSWE_RING_BC=stage_uv equivalent) |
|
|
|
`` |
CFL-only depth floor (m); see Config.h_min_cfl |
|
|
|
|
cap the depth each step (m); bounds pit blow-up on bad DEMs |
|
|
|
|
uniform infiltration rate (mm/h) on land cells of a compressed run, set either by |
|
|
|
|
shift the rainfall clock by this many seconds (driver) |
|
|
|
|
which depth maps a compressed run writes when the depth maximum is enabled: comma list of max,final |
|
|
|
|
“tif” (default): one stitched GeoTIFF per map; “shards”: each rank writes its own GeoTIFF under a VRT mosaic |
|
|
|
|
GeoTIFF writer threads |
|
|
|
|
GeoTIFF deflate level |
|
|
|
`` |
1 = delete sub-floor depth (pre-2026-08); unset = keep-h (momentum zeroed, depth kept). Both tiers bake it into their friction and forcing kernels when the module is imported, so set it before importing |
|
Changes results¶
These seven change the computed trajectory, not the speed. A run that sets one is a different numerical experiment: re-verify the case against its reference before publishing a number from it, and record the setting beside the result.
Variable |
|
Default |
Effect |
Read in |
|---|---|---|---|---|
|
|
bed-gradient limiter, one of |
|
|
|
|
1 = the L-infinity velocity norm, |
|
|
|
|
recompute the global time step every N steps of a compressed run. 1 (the default) is every step, which the application runs and the paper’s scaling series use. Above 1 the step is held for N-1 steps and shrunk by the safety factor below, which drops the per-step device-to-host read and the dt all-reduce on N-1 of every N steps; the first scaling campaign ( |
|
|
|
|
factor applied to the reused step when |
|
|
|
|
auto |
1 = ignore sub-grid channel storage in the CFL (driver). Left unset, the driver sets it to 1 when the storage floor is at least 0.20 and leaves it off below that, so the effective default depends on the case: set it explicitly for a controlled comparison |
|
|
|
AMD GPUs only: how the ROCm compiler may fuse |
|
|
|
`` |
a guard override, not a knob: 1 permits a limiter/kernel-template mismatch instead of failing loudly. The mismatch means the kernel is not applying the limiter the configuration asks for. Never for production |
|
Performance¶
These change how the work is scheduled and what is allocated, not the arithmetic:
kernel fusion, launch geometry, register caps, memory layout, communication order. The
default is the fast path except for SWE_DENSE_XY and GEOSWE_DENSE_FUSE_STEP_FORCINGS,
held at 0 so the paper’s timings reproduce out of the box and worth setting on a large run.
Two are not numerics-neutral despite sitting here: SWE_DRY_SKIP, exact only in the
configuration its row names, and SWE_DENSE_FUSE_CFL, which diverges on a full-rectangle
run. The rest are numerics-neutral by construction.
Variable |
|
Default |
Effect |
Read in |
|---|---|---|---|---|
|
|
|
stagger the dense->compressed build across N groups to bound peak memory |
|
|
|
`` |
1 = length-1 sigma placeholder when no sigma storage is used |
|
|
|
1 (default) = runs with sub-grid channel storage take the dense fused step; 0 = the split kernels. Bit-identical |
|
|
|
|
1 = the driver’s sponge, Green-Ampt/drain and CFL reduction run inside the dense fused step. Bit-identical. Four conditions must hold: |
|
|
|
|
with the previous switch: 1 (default) = also reduce the next CFL there, which the fused kernel can do only when the CFL step is sigma-free ( |
|
|
|
|
1 = with the fused step forcings, skip carried cells that cannot change. Bit-identical |
|
|
|
|
1 = the driver keeps the gathered rain field of the current frame on the device |
|
|
|
|
1 (default) = the dense solver does not allocate the unused entropic-pressure array; 0 = allocate it. Bit-identical |
|
|
|
|
1 (default) = non-blocking global dt reduction under MPI |
|
|
|
|
1 (default) = block-level CFL reduction kernel |
|
|
|
|
1 = reduce the next step’s CFL lambda inside the dense fused step, so |
|
|
|
|
1 (default) = fused residual+update on the dense path |
|
|
|
|
1 = the 2-D dense kernels map |
|
|
|
|
1 (default) = overlap the dense halo exchange with interior compute (needs an inside mask) |
|
|
|
|
register cap for the dense residual kernel: |
|
|
|
unset |
1 = skip the per-cell setup of a residual cell whose five-point depth neighbourhood is entirely dry. Injected into the kernel source at build time, and exact only where the residual is its own kernel. With the default fused steps it is refused on the compressed path ( |
|
|
|
|
1 = precompute bed gradients (default 0: computed in-kernel) |
|
|
|
|
1 = compute the CFL before the forcings stage when the fused CFL is off |
|
|
|
|
sigma-storage variant of the forcings kernel (auto) |
|
|
|
|
sigma-storage variant of the fused step (auto) |
|
|
|
|
1 (default) = fold the next-step CFL reduction into the fused compressed step |
|
|
|
|
1 = verify the fused CFL against the separate kernel every step (debug; slow) |
|
|
|
|
1 (default) = fused residual+update on the compressed path |
|
|
|
|
register cap for the compressed residual kernel: |
|
|
|
|
1 (default) = regular-neighbour fast path |
|
|
|
|
1 (default) = split regular/irregular launches |
|
|
|
|
1 = a separate residual kernel for cells whose four neighbours sit at a fixed stride, compiled into the kernel as a literal. Inert unless |
|
|
|
|
CUDA block size for the compressed residual (0 = default) |
|
|
|
|
0 (default) = skip the residual memsets |
|
|
|
|
1 (default) = one post-step kernel for rain/friction/wet-dry/depth-max, on both tiers. 0 also sends the dense path back to its split residual and update |
|
|
|
|
1 (default) = the dense forcing kernel maps threads along the contiguous array axis; 0 = the older mapping. Read when |
|
|
|
|
1 = GPU-aware MPI halo (CUDA-aware MPI on NVIDIA; Cray MPICH with |
|
|
|
|
1 (default) = one pack/unpack kernel per halo face |
|
|
|
|
1 (default) = overlap the compressed halo exchange with interior compute |
|
|
|
|
trim the CuPy memory pool every N steps |
|
|
|
`` |
reserved (research-tree rain gather). Not implemented here: |
|
|
|
|
1 (default) = evaluate the ring boundary on the GPU |
|
|
|
|
save GPU scatter tables with the cache |
|
Checked where it matters: tests/test_gpu_dense_fused_forcings.py asserts a bit-identical
final state, and that the switch engaged, for twelve of them on one GPU (five on the dense
path, seven on the compressed one); tests/mpi_bitcheck.py and
benchmark/scaling_640m/run_bitcheck.sh compare two-rank state digests across the halo
settings; and benchmark/frontier_amd/partition_check.py compares them across rank counts
on AMD GPUs. None of that runs in CI, which has no GPU (pytest -m "not gpu"; one job hands
every kernel source to nvcc, which checks that they compile, not what they compute).
Diagnostics¶
Profiling and tracing; off by default.
Variable |
|
Default |
Effect |
Read in |
|---|---|---|---|---|
|
|
unset |
1 = log device memory per stage |
|
|
unset |
1 = the dense solver prints once which step kernels a run takes (the benchmark scripts and the run driver set it) |
|
|
|
unset |
1 = skip the check that refuses the GPU backend when two CuPy builds are installed |
|
|
|
|
|
1 = count velocity-cap activations and print the total at the end of the run. It also drops the run onto the split compressed step, so do not time a run with it on |
|
|
`` |
“no_bed_source” disables the Audusse bed source (verification only) |
|
|
|
|
print dt every N steps |
|
|
|
|
per-stage timing every N steps |
|
SWE_GHOST_ETA is the research-tree alias of GEOSWE_RING_ETA (different suffix, same meaning), and SWELL_BACKEND and SWE_IGR_BACKEND are research-tree aliases of GEOSWE_BACKEND.
GEOSWE_MPI_ENV and GEOSWE_MPI_PREFIX (the benchmark MPI environment) are shell variables of the benchmark scripts and are documented there. The MPI modules also read the launcher’s rank variables (OMPI_COMM_WORLD_LOCAL_RANK, MPI_LOCALRANKID, SLURM_LOCALID) to pin each rank to a GPU, and the replay driver reads SLURM_JOB_ID and SLURM_JOB_END_TIME for its run record and deadline.