|
| 1 | +# ADMStage {#stage_adm} |
| 2 | + |
| 3 | +**Header:** `transforms/adm/adm_stage.h` |
| 4 | +**Class:** `fz::ADMStage` |
| 5 | +**Category:** Transform (lossless) |
| 6 | + |
| 7 | +**Common instantiation:** |
| 8 | +```cpp |
| 9 | +auto* adm = p.addStage<fz::ADMStage>(); |
| 10 | +adm->setDtype(fz::ADMDtype::U16); // match upstream output type |
| 11 | +``` |
| 12 | +
|
| 13 | +--- |
| 14 | +
|
| 15 | +## What it does |
| 16 | +
|
| 17 | +Remaps a `uint16_t[]` or `uint32_t[]` integer stream into a compact 8-bit symbol |
| 18 | +domain, dramatically improving the compression ratio of a downstream entropy coder |
| 19 | +(typically `ANSStage`). |
| 20 | +
|
| 21 | +- **Forward:** `uint16_t[]` or `uint32_t[]` → opaque ADM payload (`uint8_t[]`) |
| 22 | +- **Inverse:** ADM payload → original integer array (exact reconstruction) |
| 23 | +
|
| 24 | +**Algorithm:** ADM partitions the input into 512-element warp blocks and computes |
| 25 | +a per-block center (mean). Each element is then encoded as a *(code, unary-signal)* |
| 26 | +pair relative to the center. The resulting byte codes have a highly skewed, |
| 27 | +low-entropy distribution ideal for GPU rANS (`ANSStage`) or Huffman coding. |
| 28 | +
|
| 29 | +The output is an opaque ADM payload — `getOutputDataType()` returns `UNKNOWN` to |
| 30 | +opt out of downstream type checking. |
| 31 | +
|
| 32 | +--- |
| 33 | +
|
| 34 | +## Stage settings |
| 35 | +
|
| 36 | +| Setting | Type | Default | Purpose | |
| 37 | +|---|---|---|---| |
| 38 | +| `setDtype(ADMDtype)` | `ADMDtype` | `U16` | Input element type: `U16` or `U32` | |
| 39 | +
|
| 40 | +`ADMDtype::U16` expects `uint16_t` input; `ADMDtype::U32` expects `uint32_t` input. |
| 41 | +The dtype must match the upstream stage's output element type. |
| 42 | +
|
| 43 | +--- |
| 44 | +
|
| 45 | +## Typical pipeline |
| 46 | +
|
| 47 | +### Standalone (integer array input) |
| 48 | +
|
| 49 | +```cpp |
| 50 | +Pipeline p(in_bytes, MemoryStrategy::PREALLOCATE); |
| 51 | +
|
| 52 | +auto* adm = p.addStage<ADMStage>(); |
| 53 | +adm->setDtype(ADMDtype::U16); |
| 54 | +p.finalize(); |
| 55 | +
|
| 56 | +p.compress(d_in, in_bytes, stream); |
| 57 | +``` |
| 58 | + |
| 59 | +### cuSZ-style Lorenzo + ADM + ANS (recommended) |
| 60 | + |
| 61 | +```cpp |
| 62 | +Pipeline p(in_bytes, MemoryStrategy::PREALLOCATE); |
| 63 | + |
| 64 | +auto* lrz = p.addStage<LorenzoQuantStage<float, uint16_t>>(); |
| 65 | +lrz->setErrorBound(1e-3f); |
| 66 | +lrz->setQuantRadius(512); |
| 67 | +lrz->setZigzagCodes(true); |
| 68 | + |
| 69 | +auto* adm = p.addStage<ADMStage>(); |
| 70 | +adm->setDtype(ADMDtype::U16); |
| 71 | +p.connect(adm, lrz, "codes"); // connect to "codes" port |
| 72 | + |
| 73 | +auto* ans = p.addStage<ANSStage>(); |
| 74 | +p.connect(ans, adm); |
| 75 | + |
| 76 | +p.finalize(); |
| 77 | +``` |
| 78 | +
|
| 79 | +**Why ADM before ANS:** `ANSStage` operates on 8-bit symbols (256-entry alphabet). |
| 80 | +Raw `uint16_t` quantization codes have a 65536-symbol alphabet, which ANS cannot |
| 81 | +encode directly. ADM acts as a symbol-domain adapter — it remaps the wide integer |
| 82 | +stream into the 8-bit domain while preserving all information losslessly. |
| 83 | +
|
| 84 | +--- |
| 85 | +
|
| 86 | +## TOML configuration |
| 87 | +
|
| 88 | +```toml |
| 89 | +[[stage]] |
| 90 | +name = "adm" |
| 91 | +type = "ADM" |
| 92 | +dtype = "uint16" # "uint16" (default) or "uint32" |
| 93 | +inputs = [{from = "lrz", port = "codes"}] |
| 94 | +``` |
| 95 | + |
| 96 | +--- |
| 97 | + |
| 98 | +## Execution flow (CPU–GPU movement pattern) {#adm-execution} |
| 99 | + |
| 100 | +### Forward pass |
| 101 | + |
| 102 | +ADM uses one of two prefix-sum strategies depending on the number of warp blocks |
| 103 | +(`gsize = ⌈N / 512⌉`): |
| 104 | + |
| 105 | +**Decoupled look-back path** (`gsize ≤ 1024`): |
| 106 | + |
| 107 | +``` |
| 108 | +GPU ←input uint16_t[] output ADM payload uint8_t[]→ |
| 109 | + 1. adm_map_decoupled_u16 — per-warp center + encode, decoupled prefix sum |
| 110 | + 2. adm_concat_u16 — pack (code, signals) into contiguous payload |
| 111 | + └─ cudaMemcpyAsync D2H + cudaStreamSynchronize ◄── HOST BARRIER |
| 112 | + (reads actual output_lengths to determine actual_output_size_) |
| 113 | +``` |
| 114 | + |
| 115 | +**Thrust fallback path** (`gsize > 1024`, ~512 K+ elements for uint16_t): |
| 116 | + |
| 117 | +``` |
| 118 | +GPU ←input uint16_t[] output ADM payload uint8_t[]→ |
| 119 | + 1. adm_map_thrust_u16 — per-warp center + encode, Thrust exclusive_scan |
| 120 | + 2. adm_concat_u16 — pack (code, signals) into contiguous payload |
| 121 | + └─ cudaMemcpyAsync D2H + cudaStreamSynchronize ◄── HOST BARRIER |
| 122 | +``` |
| 123 | + |
| 124 | +### Inverse pass |
| 125 | + |
| 126 | +``` |
| 127 | +GPU ←input ADM payload uint8_t[] output uint16_t[]→ |
| 128 | + 1. adm_decompress_u16 — decode (code, signals) → original values using stored centers |
| 129 | +``` |
| 130 | + |
| 131 | +**Consequence:** one host barrier per compress call — the stage is not CUDA Graph |
| 132 | +compatible. `isGraphCompatible()` returns `false`. |
| 133 | + |
| 134 | +--- |
| 135 | + |
| 136 | +## Scratch buffers / device footprint |
| 137 | + |
| 138 | +`ADMStage` pre-allocates 10 persistent device buffers from the pipeline |
| 139 | +`MemoryPool` at `finalize()` time (PREALLOCATE mode) or on the first `execute()` |
| 140 | +call (MINIMAL mode). All buffers are grow-only. |
| 141 | + |
| 142 | +Let `gsize = ⌈N / 512⌉` and `sig = kMaxSignalBytes` (2 for U16, 4 for U32): |
| 143 | + |
| 144 | +| Buffer | Size formula | |
| 145 | +|---|---| |
| 146 | +| `d_signal_length_` | `gsize × 4` B | |
| 147 | +| `d_output_lengths_` | `(gsize + 1) × 4` B | |
| 148 | +| `d_centers_` | `gsize × sizeof(T)` (2 or 4 B per element) | |
| 149 | +| `d_block_flags_` | `⌈gsize × 512 / 32⌉ × 4` B | |
| 150 | +| `d_codes_` | `N` B | |
| 151 | +| `d_concat_signals_` | `N × sig` B | |
| 152 | +| `d_bit_signals_` | `N × sig` B (Thrust fallback path) | |
| 153 | +| `d_loc_offset_` | `(gsize + 1) × 4` B | |
| 154 | +| `d_prefix_state_` | `(gsize + 1) × 4` B | |
| 155 | +| `d_overflow_flag_` | `4` B (debug overflow sentinel; always allocated) | |
| 156 | + |
| 157 | +For 4096 uint16_t elements: ~68 KiB total device footprint. |
| 158 | +For 1M uint16_t elements: ~10 MiB total device footprint. |
| 159 | + |
| 160 | +--- |
| 161 | + |
| 162 | +## Serialized header |
| 163 | + |
| 164 | +The FZM stage header is 12 bytes: |
| 165 | + |
| 166 | +``` |
| 167 | +[0] dtype (0 = U16, 1 = U32) |
| 168 | +[1..3] reserved (zero) |
| 169 | +[4..11] num_elements (uint64_t LE; element count of the uncompressed input) |
| 170 | +``` |
| 171 | + |
| 172 | +`num_elements` is stored so that `estimateOutputSizes()` can return the exact |
| 173 | +decompressed byte count for inverse-pass output buffer allocation. |
| 174 | + |
| 175 | +--- |
| 176 | + |
| 177 | +## Limitations {#adm-limitations} |
| 178 | + |
| 179 | +**Not CUDA Graph compatible.** A device-to-host synchronization occurs in every |
| 180 | +forward call to read the actual payload size. `isGraphCompatible()` returns `false`. |
| 181 | + |
| 182 | +**Bounded-diff constraint.** Each GPU thread encodes `kChunk = 16` elements into a |
| 183 | +fixed `kChunk × kMaxSignalBytes` byte buffer (`32` B for U16, `64` B for U32). The |
| 184 | +number of signal bits required per element is `⌈diff / 126⌉`, where `diff` is the |
| 185 | +element's absolute deviation from the warp-block center. If the total per-thread |
| 186 | +signal bits exceed the buffer size, the kernel writes out-of-bounds. In debug |
| 187 | +builds (`NDEBUG` not defined) an `atomicOr` sentinel detects this condition and |
| 188 | +`compress_u16`/`compress_u32` throw `std::runtime_error` rather than silently |
| 189 | +corrupting data. ADM is designed for bounded quantization codes (output of |
| 190 | +`LorenzoQuantStage` or `QuantizerStage`) — arbitrary integer arrays with large |
| 191 | +value ranges are not supported. |
| 192 | + |
| 193 | +**Opaque output.** The ADM payload is not a self-describing format — it cannot be |
| 194 | +decoded without the FZM header (which stores `num_elements`). The payload must |
| 195 | +always be wrapped in a `Pipeline` with `writeToFile`/`decompressFromFile` or paired |
| 196 | +with the pipeline `decompress()` that restores the header. |
| 197 | + |
| 198 | +**Dtype must be set before `finalize()`.** `setDtype()` must be called before |
| 199 | +`pipeline.finalize()` so the correct element width is used for scratch sizing and |
| 200 | +type checking. |
| 201 | + |
| 202 | +--- |
| 203 | + |
| 204 | +## Acknowledgements |
| 205 | + |
| 206 | +`ADMStage` is a direct port of the ADM encode/decode kernels from **MANS** |
| 207 | +(Wenjing Huang, Jinwu Yang, JingKai Huang, Haoquan Long, Dingwen Tao, |
| 208 | +Guangming Tan) from `nv/adm/` in the MANS repository. Kernel logic is |
| 209 | +unchanged; changes from the original are documented at the top of |
| 210 | +`modules/transforms/adm/mapping_uint16.cu` and `mapping_uint32.cu`. |
| 211 | + |
| 212 | +> Wenjing Huang, Jinwu Yang, JingKai Huang, Haoquan Long, Dingwen Tao, |
| 213 | +> Guangming Tan. *MANS: Multidimensional Adaptive Numerical Compressor for |
| 214 | +> Scientific Data.* https://github.com/hpdps-group/MANS |
| 215 | +
|
| 216 | +See `THIRD_PARTY.md` for the full BSD-3-Clause license text. |
0 commit comments