|
| 1 | +# AdaptiveBitpackStage {#stage_adaptive_bitpack} |
| 2 | + |
| 3 | +**Header:** `modules/coders/adaptive_bitpack/adaptive_bitpack_stage.h` |
| 4 | +**Class:** `fz::AdaptiveBitpackStage<T>` |
| 5 | +**Category:** Coder (lossless) |
| 6 | + |
| 7 | +--- |
| 8 | + |
| 9 | +## What it does |
| 10 | + |
| 11 | +Per-block adaptive fixed-rate bit-plane coder — the cuSZp lossless back-end's |
| 12 | +"plain" mode, as a modular stage. It partitions the signed input into fixed-size |
| 13 | +blocks and, for each block, stores only as many bit-planes as the largest |
| 14 | +magnitude in that block requires: |
| 15 | + |
| 16 | +- **rate byte** `r` — the bit width of the largest `|value|` in the block (0 if |
| 17 | + the block is all zeros). |
| 18 | +- when `r > 0`: a **sign bitmap** (`ceil(block_size/8)` bytes) followed by `r` |
| 19 | + **bit-planes** of `ceil(block_size/8)` bytes each. Bit `j` of byte `k` of |
| 20 | + plane `p` is bit `p` of `|element 8k+j|`. |
| 21 | + |
| 22 | +Per-block payloads are concatenated using a device-wide exclusive scan of the |
| 23 | +per-block byte costs, preceded by the array of rate bytes. The archive carries no |
| 24 | +internal header of its own — `block_size` and `num_elements` live in the FZM |
| 25 | +stage header. |
| 26 | + |
| 27 | +Unlike cuSZp's fused single kernel, the offset scan here is an ordinary CUB |
| 28 | +`DeviceScan` rather than the cuSZp decoupled look-back scan: fusing predictor + |
| 29 | +quantizer + coder into one kernel (and lowering the scan to a single pass) is a |
| 30 | +job for the downstream compiler, not the stage. |
| 31 | + |
| 32 | +--- |
| 33 | + |
| 34 | +## Template parameter |
| 35 | + |
| 36 | +`T` — signed element type: `int16_t` or `int32_t` (quantizer codes or block |
| 37 | +deltas). |
| 38 | + |
| 39 | +--- |
| 40 | + |
| 41 | +## Stage settings |
| 42 | + |
| 43 | +| Setting | Purpose | Notes | |
| 44 | +|---|---|---| |
| 45 | +| `setBlockSize(n)` | Elements per block (fixed-rate granularity) | `[1, 1024]`; default 32 (cuSZp) | |
| 46 | +| `setOutlierSelection(b)` | cuSZp2 per-block plain/outlier selection | off by default; see below | |
| 47 | + |
| 48 | +```cpp |
| 49 | +auto* ab = p.addStage<AdaptiveBitpackStage<int32_t>>(); |
| 50 | +ab->setBlockSize(32); |
| 51 | +ab->setOutlierSelection(true); // cuSZp2 mode (optional) |
| 52 | +``` |
| 53 | +
|
| 54 | +### Outlier selection (cuSZp2) |
| 55 | +
|
| 56 | +With `setOutlierSelection(true)`, each block independently chooses the cheaper of |
| 57 | +two encodings: |
| 58 | +
|
| 59 | +- **plain** — pack all elements (as above), or |
| 60 | +- **outlier** — store element 0 separately as a raw 1..`sizeof(T)`-byte magnitude |
| 61 | + and pack only elements 1..n-1. |
| 62 | +
|
| 63 | +This targets non-sparse, high-smoothness data: with a block-local predictor the |
| 64 | +first element of each block is a delta-vs-0 (a full magnitude) that would inflate |
| 65 | +the whole block's bit width. Per-block metadata grows from 1 to 2 bytes |
| 66 | +(`[rate][sel]`, where `sel` bit 0 = is-outlier and bits 1-2 = outlier byte count |
| 67 | +− 1). The mode is recorded in the FZM header, so a cold decompress selects the |
| 68 | +right path automatically. This is the cuSZp **outlier** mode (SC'24); cuSZp packs |
| 69 | +the same flags into a single rate byte, which we widen to two bytes so the full |
| 70 | +`int32` rate range stays representable. |
| 71 | +
|
| 72 | +--- |
| 73 | +
|
| 74 | +## Ports |
| 75 | +
|
| 76 | +Single input → single output. |
| 77 | +
|
| 78 | +| Direction | Port | Type | |
| 79 | +|---|---|---| |
| 80 | +| Forward in / inverse out | `"output"` | `T[n]` (signed codes) | |
| 81 | +| Forward out / inverse in | `"output"` | `uint8_t[]` (archive) | |
| 82 | +
|
| 83 | +--- |
| 84 | +
|
| 85 | +## Graph compatibility |
| 86 | +
|
| 87 | +`isGraphCompatible()` is **false** — the forward path does a host-blocking D2H to |
| 88 | +read the scanned total payload length (same pattern as `BitplaneRLEStage`). |
| 89 | +
|
| 90 | +--- |
| 91 | +
|
| 92 | +## Typical pipeline (cuSZp-style) |
| 93 | +
|
| 94 | +```cpp |
| 95 | +auto* quant = p.addStage<QuantizerStage<float, uint32_t>>(); |
| 96 | +quant->setErrorBound(1e-3f); |
| 97 | +quant->setErrorBoundMode(ErrorBoundMode::ABS); |
| 98 | +quant->setLinearMode(true); // signed INT32 codes, no outliers |
| 99 | +
|
| 100 | +auto* lrz = p.addStage<LorenzoStage<int32_t>>(); |
| 101 | +lrz->setBlockSize(32); // block-local 1-D delta |
| 102 | +
|
| 103 | +auto* ab = p.addStage<AdaptiveBitpackStage<int32_t>>(); |
| 104 | +ab->setBlockSize(32); |
| 105 | +
|
| 106 | +p.connect(lrz, quant, "codes"); |
| 107 | +p.connect(ab, lrz); |
| 108 | +p.finalize(); |
| 109 | +``` |
| 110 | + |
| 111 | +`block_size` on the coder need not match the Lorenzo block, but for faithful |
| 112 | +cuSZp both are 32. |
| 113 | + |
| 114 | +--- |
| 115 | + |
| 116 | +## TOML |
| 117 | + |
| 118 | +```toml |
| 119 | +[[stage]] |
| 120 | +type = "AdaptiveBitpack" |
| 121 | +input_type = "int32" # or "int16" |
| 122 | +block_size = 32 |
| 123 | +outlier_selection = false # true = cuSZp2 per-block plain/outlier selection |
| 124 | +``` |
| 125 | + |
| 126 | +--- |
| 127 | + |
| 128 | +## Acknowledgements |
| 129 | + |
| 130 | +The per-block adaptive fixed-rate bit-plane scheme is from cuSZp (Yafan Huang et |
| 131 | +al., SC'23/SC'24). This is an independent reimplementation of the published |
| 132 | +scheme — no cuSZp source is copied. See `THIRD_PARTY.md` and |
| 133 | +`memory/cuszp_stages.md`. |
0 commit comments