Skip to content

Commit 7cb1577

Browse files
committed
adds adaptive bitpacking kernels from cuszp2
1 parent 6f7c11a commit 7cb1577

15 files changed

Lines changed: 1494 additions & 0 deletions

‎CHANGELOG.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -17,6 +17,7 @@ Version numbers follow [Semantic Versioning](https://semver.org/).
1717
- `tests/helpers/stage_harness.h` `pipeline_round_trip<T>()` was unconditionally calling `cudaFree(d_dec)` on the decompress output. When the pipeline uses `pool_managed_decomp_=true` (the default), `d_dec` is owned by `MemoryPool::graph_allocations_` and freed by `Pipeline::~Pipeline()` via `cudaFreeAsync`. The harness's extra `cudaFree` was a double-free that surfaced under compute-sanitizer as `cudaErrorInvalidValue` from `cudaFreeAsync` at teardown (one error per affected test), and the sticky error propagated to the next test's `cudaGetLastError()` checks in `CompressionDAG::addStage`. Harness now branches on `pipeline.isPoolManagedDecompOutput()` — only `cudaFree`s when the caller opted into caller-owned mode via `setPoolManagedDecompOutput(false)`. `pipeline_file_round_trip<T>()` unchanged (it uses the static `Pipeline::decompressFromFile`, which always returns a `cudaMalloc`-owned pointer). Also fixed two callers in `tests/stages/test_ginterp.cpp` (GI27, GI29) that replicated the same pattern inline. Result: `test_ginterp`, `test_lorenzo`, `test_quantizer`, `test_bitpack`, `test_ans`, `test_graph_capture` all report `ERROR SUMMARY: 0 errors` under `compute-sanitizer --tool=memcheck` (was 1+ pre-existing errors on most). `test_huffman` still has a pre-existing kernel-side issue unrelated to ownership.
1818

1919
### Added
20+
- `AdaptiveBitpackStage<T>` (`modules/coders/adaptive_bitpack/`, `StageType::ADAPTIVE_BITPACK = 24`) — a per-block adaptive fixed-rate bit-plane coder: the cuSZp lossless back-end's "plain" mode as a modular stage. For each block of `setBlockSize(n)` signed elements (`int16_t`/`int32_t`; default n=32, range [1,1024]) it emits a rate byte (bit width of the block's largest magnitude) and, when non-zero, a sign bitmap + that many bit-planes (`ceil(n/8)` bytes each). Per-block payload offsets come from an ordinary CUB `DeviceScan::ExclusiveSum` over per-block byte costs — **not** cuSZp's fused decoupled look-back scan, which is a fusion-time lowering left to the downstream compiler. Single signed-codes input → single `uint8` archive (rate region + scan-packed payloads); `block_size`+`num_elements` live in the 16-byte FZM header (the archive needs no internal header). Kernels are one-thread-per-block for verifiability; **independent reimplementation — no cuSZp source copied** (THIRD_PARTY cuSZp section added). Not graph-compatible (forward D2H for the scanned length, like `BitplaneRLE`). Completes the modular cuSZp front-to-back pipeline `QuantizerStage(linear) → LorenzoStage(setBlockSize) → AdaptiveBitpackStage`. Also supports cuSZp2's per-block **outlier selection** via `setOutlierSelection(true)`: each block independently picks the cheaper of plain packing vs. extracting element 0 as a raw 1..`sizeof(T)`-byte outlier and packing only the rest (targets non-sparse data where the block-local predictor's first element is a full-magnitude delta-vs-0). Per-block metadata widens from 1 to 2 bytes (`[rate][sel]`) so the full `int32` rate range stays representable — cuSZp bit-packs the same flags into one byte; we keep them separate. The mode is carried in the FZM header. Registered in `stage_factory.h`, `fzgpumodules.h`, and TOML (`type = "AdaptiveBitpack"`, `input_type`, `block_size`, `outlier_selection`). `tests/stages/test_adaptive_bitpack.cpp` — 19 tests (int16/int32 round-trips, all-zero blocks, every fixed-rate 0–31, partial final block, INT_MIN/sign extremes, header round-trip, `setBlockSize` bounds, port/type contract, graph-compat=false, compression-ratio sanity, the full `Quantizer(linear)→Lorenzo(block=32)→AdaptiveBitpack` float pipeline, and outlier-mode round-trips covering every outlier byte-count, outlier-beats-plain, partial/single blocks, INT_MIN outlier, header flag, and the outlier-mode pipeline); all pass and compute-sanitizer-clean (0 errors). Docs: `docs/stages/adaptive_bitpack.md`; design notes in `memory/cuszp_stages.md` §7.3–7.4.
2021
- `LorenzoStage::setBlockSize(uint32_t)` — an explicit, launch-decoupled 1-D block-local reset period (cuSZp-style; cuSZp uses 32). `n=0` (default) keeps the existing behavior (N-D inclusion-exclusion delta selected by `dims_`; the 1-D path already reset per 256-thread launch block). `n>0` forces the 1-D path over the flattened array and restarts the prediction chain every `n` elements regardless of launch config or `dims_`. Implemented by launching the existing `lorenzo_delta_1d_kernel` / `lorenzo_scan_1d_kernel` with `blockDim == n` (both already reset per CUDA block), so no new kernels — the 1-D launchers gained a trailing `unsigned block_threads = 256` parameter. `n` must be in `[1, 1024]` (`setBlockSize` throws otherwise; the inverse scans one segment per CUDA block of `n` threads). Graph-compatible. New `block_size` field in `LorenzoConfig` (struct 16→20 B; `deserializeHeader` accepts legacy 16-byte headers and defaults it to 0). `tests/stages/test_lorenzo.cpp` — added LZ10–LZ14 (block round-trip, ragged final block, header round-trip incl. legacy, `setBlockSize` bounds, and a `Quantizer(linear)→Lorenzo(block=32)` cuSZp front-end pipeline tying together §7.1+§7.2); suite 14/14. Docs: `docs/stages/lorenzo.md`. See `memory/cuszp_stages.md` §7.2.
2122
- `QuantizerStage::setLinearMode(bool)` — a cuSZp-style linear / no-outlier quantization mode (ABS/NOA only). Maps each value to `q = round(x / (2·eb))` stored two's-complement in `TCode` with **no radius clamp, no outlier ports, no zigzag**; the forward kernel `quantizer_linear_fwd_kernel` is a pure memory-bound map (the regular forward kernel's only `atomicAdd` and only divergent branch — the outlier path — are removed), and the inverse reuses `quantizer_abs_inv_kernel<…,false>` with no memset/scatter. Single `"codes"` output (forward) / single input (inverse); the codes are *declared* as the signed DataType (`UINT16→INT16`, `UINT32→INT32`) so they connect directly to `LorenzoStage<intN>` — no new signed-`TCode` instantiations needed. A bin overflowing `TCode` wraps (no outlier fallback), so size `TCode` wide enough (use `uint32_t`). Fully graph-compatible in ABS mode (no D2H, no count readback). Mutually exclusive with REL / inplace / zigzag — each throws at the first `compress()`. New `linear_mode` byte in `QuantizerConfig` (reuses former `_pad[0]`; struct size unchanged at 36 B, old headers default it to 0). Intended front-end for the modular cuSZp pipeline `Quantizer(linear) → LorenzoStage(blockSize=32) → AdaptiveBitpackStage` (see `memory/cuszp_stages.md` §7.1). `tests/stages/test_quantizer.cpp` — added QuantizerLinear suite (5 tests: ABS/NOA round-trips, port/type contract, header round-trip, incompatible-mode throws); full suite 34/34. Docs: `docs/stages/quantizer.md`.
2223
- `fzgmod-cli --report-json <file>` — writes a standalone, machine-readable JSON report for the benchmarking harness to parse (compress/decompress/benchmark modes). New `src/utils/cli/report_json.{h,cpp}` with a plain-data `ReportData` model decoupled from CLI internals; the writer is the single source of truth for derived values (ratio, bitrate, throughput, median/min/max) while raw counts (`original_bytes`, `compressed_bytes`, `num_elements`) are always emitted so the harness can recompute. Schema (`schema_version` 1.0): `tool` (name/version/git_sha, injected via CMake `FZGMOD_VERSION`/`FZGMOD_GIT_SHA` defines), `status`+`error_message` (a forced failure still emits valid JSON with a non-zero exit, via `emit_error_report` on the top-level catch), `config` echo, `size`, `timing` (per-rep `device_ms`/`host_wall_ms` arrays with `timing_method: cuda_events_dag`, labeled honestly now that `dag_elapsed_ms` is event-based), derived `throughput` (device-median basis), `memory`, optional `quality` (when an original is available), and per-stage `stages[]`. `--report-json` forces profiling on (without enabling the stdout perf table) so device timing + stage breakdown are populated even for single-shot compress/decompress. `tests/test_cli.cpp` — added `ReportJsonCompressEmitsValidFile` (asserts schema/status/raw-count fields + stage block) and `ReportJsonErrorStillEmitsValidJsonAndNonZeroExit`; full CLI suite 7/7. Spec: `memory/report_json_spec.md`.

‎CMakeLists.txt‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -243,6 +243,8 @@ add_library(fzgmod_modules
243243
modules/coders/rle/rle.cu
244244
modules/coders/rze/rze_stage.cu
245245
modules/coders/bitpack/bitpack_stage.cu
246+
modules/coders/adaptive_bitpack/adaptive_bitpack_kernels.cu
247+
modules/coders/adaptive_bitpack/adaptive_bitpack_stage.cu
246248
modules/shufflers/bitshuffle/bitshuffle_stage.cu
247249
modules/fused/lorenzo_quant/lorenzo_quant.cu
248250
modules/fused/lorenzo_quant/lorenzo_quant_nd.cu

‎THIRD_PARTY.md‎

Lines changed: 29 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -164,6 +164,35 @@ the [cuSZ / PHF](#cusz--phf) section above.
164164

165165
---
166166

167+
## cuSZp
168+
169+
**Used by:** `AdaptiveBitpackStage` (`modules/coders/adaptive_bitpack/`),
170+
and the `linear` mode of `QuantizerStage` + the `setBlockSize` option of
171+
`LorenzoStage`.
172+
173+
**Relationship:**
174+
- These are **independent reimplementations of published cuSZp schemes — no
175+
cuSZp source code is copied.** `AdaptiveBitpackStage` implements cuSZp's
176+
per-block adaptive fixed-rate bit-plane "plain" encoding with a byte-granular
177+
layout, one-thread-per-block kernels, and an ordinary CUB `DeviceScan` for the
178+
per-block offsets (cuSZp uses a fused single kernel with a decoupled look-back
179+
scan; that fusion is left to a downstream compiler). `QuantizerStage`'s linear
180+
mode reproduces cuSZp's `q = round(x / 2·eb)` mapping with no radius/outlier
181+
fallback, and `LorenzoStage::setBlockSize` reproduces cuSZp's block-local 1-D
182+
delta. The reference codebase (for cross-checking only, not vendored) lives at
183+
`compressors/cuSZp2/`; design notes are in `memory/cuszp_stages.md`.
184+
185+
Original authors: Yafan Huang et al.
186+
Papers: Yafan Huang, Sheng Di, Xiaodong Yu, Guanpeng Li, Franck Cappello,
187+
"cuSZp: An Ultra-fast GPU Error-bounded Lossy Compression Framework with
188+
Optimized End-to-End Performance", SC '23; and "cuSZp2: A GPU Lossy Compressor
189+
with Extreme Throughput and Optimized Compression Ratio", SC '24.
190+
191+
**License:** cuSZp is BSD-3-Clause (Argonne National Laboratory / University of
192+
Iowa). As no source is copied, this is an algorithmic attribution.
193+
194+
---
195+
167196
## cuSZ-Hi
168197

169198
**Used by:** `GInterpStage` (`modules/fused/ginterp/`)

‎docs/stages/adaptive_bitpack.md‎

Lines changed: 133 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,133 @@
1+
# AdaptiveBitpackStage {#stage_adaptive_bitpack}
2+
3+
**Header:** `modules/coders/adaptive_bitpack/adaptive_bitpack_stage.h`
4+
**Class:** `fz::AdaptiveBitpackStage<T>`
5+
**Category:** Coder (lossless)
6+
7+
---
8+
9+
## What it does
10+
11+
Per-block adaptive fixed-rate bit-plane coder — the cuSZp lossless back-end's
12+
"plain" mode, as a modular stage. It partitions the signed input into fixed-size
13+
blocks and, for each block, stores only as many bit-planes as the largest
14+
magnitude in that block requires:
15+
16+
- **rate byte** `r` — the bit width of the largest `|value|` in the block (0 if
17+
the block is all zeros).
18+
- when `r > 0`: a **sign bitmap** (`ceil(block_size/8)` bytes) followed by `r`
19+
**bit-planes** of `ceil(block_size/8)` bytes each. Bit `j` of byte `k` of
20+
plane `p` is bit `p` of `|element 8k+j|`.
21+
22+
Per-block payloads are concatenated using a device-wide exclusive scan of the
23+
per-block byte costs, preceded by the array of rate bytes. The archive carries no
24+
internal header of its own — `block_size` and `num_elements` live in the FZM
25+
stage header.
26+
27+
Unlike cuSZp's fused single kernel, the offset scan here is an ordinary CUB
28+
`DeviceScan` rather than the cuSZp decoupled look-back scan: fusing predictor +
29+
quantizer + coder into one kernel (and lowering the scan to a single pass) is a
30+
job for the downstream compiler, not the stage.
31+
32+
---
33+
34+
## Template parameter
35+
36+
`T` — signed element type: `int16_t` or `int32_t` (quantizer codes or block
37+
deltas).
38+
39+
---
40+
41+
## Stage settings
42+
43+
| Setting | Purpose | Notes |
44+
|---|---|---|
45+
| `setBlockSize(n)` | Elements per block (fixed-rate granularity) | `[1, 1024]`; default 32 (cuSZp) |
46+
| `setOutlierSelection(b)` | cuSZp2 per-block plain/outlier selection | off by default; see below |
47+
48+
```cpp
49+
auto* ab = p.addStage<AdaptiveBitpackStage<int32_t>>();
50+
ab->setBlockSize(32);
51+
ab->setOutlierSelection(true); // cuSZp2 mode (optional)
52+
```
53+
54+
### Outlier selection (cuSZp2)
55+
56+
With `setOutlierSelection(true)`, each block independently chooses the cheaper of
57+
two encodings:
58+
59+
- **plain** — pack all elements (as above), or
60+
- **outlier** — store element 0 separately as a raw 1..`sizeof(T)`-byte magnitude
61+
and pack only elements 1..n-1.
62+
63+
This targets non-sparse, high-smoothness data: with a block-local predictor the
64+
first element of each block is a delta-vs-0 (a full magnitude) that would inflate
65+
the whole block's bit width. Per-block metadata grows from 1 to 2 bytes
66+
(`[rate][sel]`, where `sel` bit 0 = is-outlier and bits 1-2 = outlier byte count
67+
− 1). The mode is recorded in the FZM header, so a cold decompress selects the
68+
right path automatically. This is the cuSZp **outlier** mode (SC'24); cuSZp packs
69+
the same flags into a single rate byte, which we widen to two bytes so the full
70+
`int32` rate range stays representable.
71+
72+
---
73+
74+
## Ports
75+
76+
Single input → single output.
77+
78+
| Direction | Port | Type |
79+
|---|---|---|
80+
| Forward in / inverse out | `"output"` | `T[n]` (signed codes) |
81+
| Forward out / inverse in | `"output"` | `uint8_t[]` (archive) |
82+
83+
---
84+
85+
## Graph compatibility
86+
87+
`isGraphCompatible()` is **false** — the forward path does a host-blocking D2H to
88+
read the scanned total payload length (same pattern as `BitplaneRLEStage`).
89+
90+
---
91+
92+
## Typical pipeline (cuSZp-style)
93+
94+
```cpp
95+
auto* quant = p.addStage<QuantizerStage<float, uint32_t>>();
96+
quant->setErrorBound(1e-3f);
97+
quant->setErrorBoundMode(ErrorBoundMode::ABS);
98+
quant->setLinearMode(true); // signed INT32 codes, no outliers
99+
100+
auto* lrz = p.addStage<LorenzoStage<int32_t>>();
101+
lrz->setBlockSize(32); // block-local 1-D delta
102+
103+
auto* ab = p.addStage<AdaptiveBitpackStage<int32_t>>();
104+
ab->setBlockSize(32);
105+
106+
p.connect(lrz, quant, "codes");
107+
p.connect(ab, lrz);
108+
p.finalize();
109+
```
110+
111+
`block_size` on the coder need not match the Lorenzo block, but for faithful
112+
cuSZp both are 32.
113+
114+
---
115+
116+
## TOML
117+
118+
```toml
119+
[[stage]]
120+
type = "AdaptiveBitpack"
121+
input_type = "int32" # or "int16"
122+
block_size = 32
123+
outlier_selection = false # true = cuSZp2 per-block plain/outlier selection
124+
```
125+
126+
---
127+
128+
## Acknowledgements
129+
130+
The per-block adaptive fixed-rate bit-plane scheme is from cuSZp (Yafan Huang et
131+
al., SC'23/SC'24). This is an independent reimplementation of the published
132+
scheme — no cuSZp source is copied. See `THIRD_PARTY.md` and
133+
`memory/cuszp_stages.md`.

‎docs/stages/coders.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -7,3 +7,4 @@
77
| \subpage stage_rle | Run-length encoding |
88
| \subpage stage_rze | Recursive zero-byte elimination |
99
| \subpage stage_bitpack | Dense bit-packing of fixed-width integers |
10+
| \subpage stage_adaptive_bitpack | Per-block adaptive fixed-rate bit-plane coding (cuSZp plain mode) |

‎include/fzgpumodules.h‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -22,6 +22,7 @@
2222
#include "predictors/lorenzo/lorenzo_stage.h"
2323
#include "coders/rle/rle.h"
2424
#include "coders/bitpack/bitpack_stage.h"
25+
#include "coders/adaptive_bitpack/adaptive_bitpack_stage.h"
2526
#include "coders/rze/rze_stage.h"
2627
#include "fused/lorenzo_quant/lorenzo_quant.h"
2728
#include "quantizers/quantizer/quantizer.h"

‎include/fzm_format.h‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -94,6 +94,7 @@ enum class StageType : uint16_t {
9494
ADM = 19, ///< Adaptive Data Mapping transform (MANS)
9595
G_INTERP = 22, ///< Spline interpolation predictor + quantizer (cuSZ-Hi G-Interp)
9696
BITPLANE_RLE = 23, ///< Fused bitplane transpose + zero-byte RLE (FZ-GPU lossless encoder)
97+
ADAPTIVE_BITPACK = 24, ///< Per-block adaptive fixed-rate bit-plane coder (cuSZp plain mode)
9798
};
9899

99100
/**
@@ -319,6 +320,7 @@ inline std::string stageTypeToString(StageType type) {
319320
case StageType::ADM: return "ADM";
320321
case StageType::G_INTERP: return "GInterp";
321322
case StageType::BITPLANE_RLE: return "BitplaneRLE";
323+
case StageType::ADAPTIVE_BITPACK: return "AdaptiveBitpack";
322324
default: return "Unknown";
323325
}
324326
}

‎include/stage/stage_factory.h‎

Lines changed: 15 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -16,6 +16,7 @@
1616
#include "shufflers/bitshuffle/bitshuffle_stage.h"
1717
#include "coders/rze/rze_stage.h"
1818
#include "coders/bitpack/bitpack_stage.h"
19+
#include "coders/adaptive_bitpack/adaptive_bitpack_stage.h"
1920
#include "coders/huffman/huffman_stage.h"
2021
#include "coders/ans/ans_stage.h"
2122
#include "transforms/adm/adm_stage.h"
@@ -297,6 +298,20 @@ inline Stage* createStage(StageType type, const uint8_t* config, size_t config_s
297298
break;
298299
}
299300

301+
case StageType::ADAPTIVE_BITPACK: {
302+
// config[0] holds the DataType of T (INT16 / INT32).
303+
DataType dt = (config_size > 0)
304+
? static_cast<DataType>(config[0])
305+
: DataType::INT32;
306+
if (dt == DataType::INT16) stage = new AdaptiveBitpackStage<int16_t>();
307+
else if (dt == DataType::INT32) stage = new AdaptiveBitpackStage<int32_t>();
308+
else throw std::runtime_error(
309+
"Unsupported AdaptiveBitpackStage DataType: "
310+
+ std::to_string(static_cast<int>(dt)));
311+
stage->deserializeHeader(config, config_size);
312+
break;
313+
}
314+
300315
default:
301316
throw std::runtime_error("Unknown stage type: "
302317
+ std::to_string(static_cast<uint16_t>(type)));

0 commit comments

Comments
 (0)