You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit b3c2c25
Browse filesBrowse the repository at this point in the historyBrowse files
Copy file name to clipboardExpand all lines: include/backend/atomics.h
+41-14Lines changed: 41 additions & 14 deletions
Original file line number
Diff line number
Diff line change
@@ -12,27 +12,52 @@
12
12
* known to be in the same block, so scoping the atomic down from "device"
13
13
* saves the extra memory-fence cost of a full device-wide atomic.
14
14
*
15
-
* ROCm's HIP declares `atomicAdd_block`/`atomicOr_block` with identical names
16
-
* and signatures to CUDA's (`hip/amd_detail/amd_hip_atomic.h`) — unlike the
17
-
* warp shuffle/ballot family (see warp.h), there is no warp/wavefront-width
18
-
* dependency in a block-scope atomic, so naively this should "just hipify".
19
-
* That said, this codebase has been burned before by CUDA intrinsics that
20
-
* looked source-portable but silently diverged under HIP (see warp.h's
21
-
* `__shfl_down_sync` width story), and — unlike every other intrinsic family
22
-
* touched during the HIP port — these have **not yet been exercised or
23
-
* verified on real AMD hardware** (no existing call site in this codebase
24
-
* uses any block-scoped atomic). This header exists so there is exactly one
25
-
* place to patch if that verification turns up a divergence, rather than
26
-
* scattering raw `atomicAdd_block`/`atomicOr_block` calls through the RARE/RAZE
27
-
* kernels. Treat the HIP branch below as "expected to work, not yet confirmed
28
-
* on hardware" until it has been built and run on the project's MI100 target.
15
+
* The `_block` suffix family is CUDA-only spelling. Contrary to this header's
16
+
* original expectation ("ROCm declares them with identical names"), ROCm 6.4.1
17
+
* declares neither `atomicAdd_block` nor `atomicOr_block` anywhere under
18
+
* `hip/` — verified by grep against the installed toolchain, and by the
19
+
* `use of undeclared identifier 'atomicAdd_block'` error the first HIP build
20
+
* of the RARE/RAZE stages produced. This is exactly the divergence the header
21
+
* was created to absorb, so the patch lands here and the call sites are
22
+
* untouched.
23
+
*
24
+
* HIP's equivalent is the scoped-atomic builtin family
25
+
* (`__hip_atomic_fetch_add`/`_or` + `__HIP_MEMORY_SCOPE_WORKGROUP`), where
26
+
* "workgroup" is AMD's name for a CUDA block. Memory ordering is
27
+
* `__ATOMIC_RELAXED` to match CUDA's `atomic*_block`, which carry no ordering
28
+
* guarantees beyond the atomicity of the read-modify-write itself — the RARE/
29
+
* RAZE call sites (per-block histogram accumulation; OR-ing a straddling
30
+
* partial word into shared memory) separately `__syncthreads()` before reading
31
+
* the accumulated result, so they rely on that barrier for ordering, not on
32
+
* the atomic.
33
+
*
34
+
* Unlike the warp shuffle/ballot family (see warp.h), a block-scoped atomic has
35
+
* no wavefront-width dependency, so there is no 32-vs-64-lane subtlety here.
29
36
*/
30
37
31
38
#include"backend/api.h"
32
39
33
40
namespacefz {
34
41
namespacebackend {
35
42
43
+
#if defined(FZGMOD_BACKEND_HIP)
44
+
45
+
/** Block-scoped atomic add. `T` must be an arithmetic type the HIP scoped-atomic builtin accepts (int, unsigned, unsigned long long, float, double). */
46
+
template <typename T>
47
+
__device__ inline T atomicAddBlock(T* addr, T val) {
0 commit comments