Skip to content

Commit f5b2f10

Browse files
committed
Add strict high-precision linear quantizer arithmetic and power-of-two bound policy
linear_high_precision keeps both coordinates and the user bound in double, rounds strict f32 metadata downward, and reserves reconstruction rounding so a tight requested bound is no longer violated. Optional power_of_two_bound rounds the resolved ABS/NOA/PREL half-bound down to a power of two. Both flags persist in the existing 72-byte header, round-trip through TOML, and stay opt-in. Adds the fzgmod-profile-quantizer-linear four-way comparison benchmark.
1 parent bda3014 commit f5b2f10

11 files changed

Lines changed: 572 additions & 56 deletions

File tree

‎CHANGELOG.md‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -10,6 +10,7 @@ Version numbers follow [Semantic Versioning](https://semver.org/).
1010
## [Unreleased] — 2.0.0
1111

1212
### Added
13+
- **Strict linear-Quantizer arithmetic and an optional SLEEK-style power-of-two bound policy.** `linear_high_precision` uses a precomputed double reciprocal on encode, direct integer-to-double reconstruction on decode, and an internal half-ULP reserve for arbitrary steps; `power_of_two_bound` rounds the resolved ABS/NOA/PREL half-bound downward (1x–2x tighter). Both flags persist in the existing 72-byte header by using previously-zero padding, round-trip through TOML, and remain opt-in. Added `fzgmod-profile-quantizer-linear` for the four-way float/double/power2 comparison.
1314
- **Explicit-ownership execution API on `Pipeline` — ownership is now expressed by the return type instead of by `void**` vs `void*` and by mutable pipeline state.** New backend-neutral value types in `include/pipeline/device_buffer.h` (`DeviceSpan`, `ConstDeviceSpan`, non-owning `BorrowedDeviceBuffer`, move-only `OwnedDeviceBuffer` that records its device and frees through the backend facade, never a hard-coded `cudaFree`), plus span wrappers `compress()`, `compressInto()`, `decompressBorrowed()`, `decompressOwned()`, `decompressInto()`, and `decompressIntoAsync()`. They delegate to the existing pointer overloads, so behavior is unchanged and nothing is deprecated; `decompressBorrowed()`/`decompressOwned()` deliberately ignore `setPoolManagedDecompOutput()` and restore the flag. Tests: `tests/pipeline/test_device_buffer.cpp` (8 cases: move-only/non-owning traits, borrow round-trip, byte-identical `compressInto`, owned-buffer reclamation, ownership independent of the pipeline flag, capacity failure, move/release). 53/53 ctest. Step 1-2 of `memory/public_api_evolution.md`.
1415
- Added `fzgmod-profile-huffman-throughput`, a direct-stage HostCoordinated/DeviceResident latency and throughput benchmark with cold/warm codebook cases, CUDA-event and host timing, multiple sizes/distributions, round-trip checks, deterministic archive comparison, and CSV output.
1516
- Added `HuffmanExecutionMode::DeviceResident`: cuSZ-compatible canonical tree/forward/reverse-book construction directly from device histograms, device-side partition scan and PHF assembly for PerBlock/Adaptive/Fixed books, safe range validation, device-side header parsing on decode, terminal fixed-book CUDA Graph support, and automatic exact-size readback when Huffman feeds a downstream stage.
@@ -21,6 +22,7 @@ Version numbers follow [Semantic Versioning](https://semver.org/).
2122
- Removed Huffman's experimental `Fine` encode mode, its public API and diagnostics, TOML/card option, documentation, profiling selection, and mode-specific tests; `HuffmanStage` now exposes only the supported cuSZ coarse-grained encoder.
2223

2324
### Fixed
25+
- **Tight linear quantization could still violate the bound after fixing integer overflow.** Float reciprocal multiplication selected adjacent bins, final f32 reconstruction consumed additional error, the inverse first cast signed codes to float—dropping low bits above 2^24—and the runtime user bound was narrowed to float before an f64 range multiplication. The strict path now keeps both coordinate directions and the user-bound parameter in double, rounds strict f32 metadata downward rather than loosening it, and reserves reconstruction rounding. A captured CESM FLNTC `rel_range=1e-7` case changed from a 1.744x miss to bit-exact reconstruction; an 18-cell CESM/HURR/NYX smoke matrix passes every requested bound.
2426
- **Linear quantization silently wrapped bins outside signed `TCode` range.** Tight NOA bounds on offset-valued fields can require indices much wider than int32 (S3D `N2` at `1e-6` needs roughly 3.3e10). The linear kernel now uses a guarded int64 rounding intermediate, records any signed-code overflow in device scratch, and `postStreamSync()` refuses the run with remediation guidance instead of producing a corrupt archive. Regression test: `QuantizerLinear.RefusesCodeOverflow`; rationale: `CN-QUANT-3`.
2527
- Docker image builds now retain `THIRD_PARTY.md` in the build context so the CMake install step can package the required third-party notices.
2628
- Installed packages now include the project `LICENSE` and `THIRD_PARTY.md`; synchronized all 30-stage documentation/catalog surfaces, completed newer-stage attribution, and corrected the cuSZ-Hi/SZx/ROIBIN-SZ license records.

‎docs/cards/fzgm-quantizer.json‎

Lines changed: 12 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -60,6 +60,18 @@
6060
"type": "bool",
6161
"default": false,
6262
"description": "ABS/NOA: linear/no-outlier mode (cuSZp-style). Emits raw signed codes q=round(x/2·eb), no radius clamp, no outlier ports. Pair with LorenzoStage(setBlockSize) → AdaptiveBitpackStage."
63+
},
64+
{
65+
"name": "linearHighPrecision",
66+
"type": "bool",
67+
"default": false,
68+
"description": "Linear mode only: double coordinate and reconstruction arithmetic plus an internal rounding reserve for a strict TInput error bound."
69+
},
70+
{
71+
"name": "powerOfTwoBound",
72+
"type": "bool",
73+
"default": false,
74+
"description": "ABS/NOA/PREL: round the resolved absolute half-bound downward to a power of two (1x to 2x tighter than requested), following SLEEK's bound policy."
6375
}
6476
],
6577

‎docs/config_file.md‎

Lines changed: 7 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -171,6 +171,13 @@ example `setBlockSize(32)` is `block_size = 32`).
171171
| `SZp` | `SZpStage` | \ref stage_szp "SZpStage" |
172172
| `Merge` | `MergeStage` | \ref stage_merge "MergeStage" |
173173
| `ROIBinSplit` | `ROIBinSplitStage` | \ref stage_roibin_split "ROIBinSplitStage" |
174+
175+
For a cuSZp-style `Quantizer` with a strict requested bound, set
176+
`linear_mode = true` and `linear_high_precision = true`. The optional
177+
`power_of_two_bound = true` rounds the resolved ABS/NOA/PREL absolute bound
178+
downward to a power of two; it is therefore a tighter, separately labelled
179+
rate-distortion configuration. See the Quantizer reference for constraints and
180+
effective-bound semantics.
174181
---
175182

176183
## Complete Examples

‎docs/stages/quantizer.md‎

Lines changed: 57 additions & 11 deletions
Original file line numberDiff line numberDiff line change
@@ -47,6 +47,8 @@ Using any other combination will result in a linker error. Most common: `Quantiz
4747
| `setOutlierThreshold(t)` | Force outliers | ABS/NOA only; `abs(x) >= t` -> outlier |
4848
| `setInplaceOutliers(enable)` | Embed outliers in codes | ABS/NOA only; see constraints below |
4949
| `setLinearMode(enable)` | Signed codes, no outliers | ABS/NOA only; cuSZp-style; see below |
50+
| `setLinearHighPrecision(enable)` | Strict linear coordinate arithmetic | Linear mode only; double forward/inverse arithmetic and rounding reserve |
51+
| `setPowerOfTwoBound(enable)` | Tighten uniform bound to a power of two | ABS/NOA/PREL; rounds down, so effective EB is 1x–2x tighter |
5052
| `setValueBase(v)` | Precomputed value range | NOA only; optional, see below |
5153
| `setDither(enable)` | Dithered ("_R"-style) reconstruction | ABS/NOA/REL; incompatible with linear/inplace modes; see below |
5254
| `setDitherSeed(seed)` | Seed for the dither offset | Persisted in the header; default 0 |
@@ -104,22 +106,22 @@ outlier mechanism at all**.
104106

105107
---
106108

107-
## Linear / no-outlier mode (ABS/NOA only)
109+
## Linear / no-outlier mode (ABS/NOA/PREL only)
108110

109111
`setLinearMode(true)` selects a cuSZp-style quantizer: each value is mapped to
110112
`q = round(x / (2 · eb))` and stored as two's-complement in `TCode`, with **no
111113
radius clamp, no outlier ports, no zigzag**. The forward kernel is a pure
112-
memory-bound map — the only atomic and the only divergent branch of the regular
113-
forward kernel (the outlier path) are removed — so it is strictly faster and is
114-
fully **graph-compatible** in ABS mode (no D2H, no outlier-count readback).
114+
memory-bound map: the outlier atomic and divergent branch of the regular forward
115+
kernel are removed. A four-byte overflow flag is read after execution, so linear
116+
mode is not CUDA-Graph compatible.
115117

116-
Because there is no outlier fallback, a bin that overflows `TCode` simply wraps;
117-
**size TCode wide enough for the data** (use `uint32_t`). The codes are
118+
Because there is no outlier fallback, a bin outside signed `TCode` range is a
119+
hard error; **size TCode wide enough for the data** (use `uint32_t`). The codes are
118120
*declared* by the stage as the **signed** DataType (`UINT16→INT16`,
119121
`UINT32→INT32`) so they connect directly to a downstream `LorenzoStage<intN>`.
120122

121123
Constraints (each throws at the first `compress()` if violated):
122-
- Valid only with `ABS` and `NOA` error-bound modes (not `REL`).
124+
- Valid only with uniform `ABS`, `NOA`, and `PREL` error-bound modes (not `REL`).
123125
- Mutually exclusive with `setInplaceOutliers(true)` and `setZigzagCodes(true)`.
124126

125127
Intended front-end for the cuSZp-style modular pipeline:
@@ -129,11 +131,52 @@ auto* quant = p.addStage<QuantizerStage<float, uint32_t>>();
129131
quant->setErrorBound(1e-3f);
130132
quant->setErrorBoundMode(ErrorBoundMode::ABS);
131133
quant->setLinearMode(true); // signed INT32 codes, no outliers
134+
quant->setLinearHighPrecision(true); // strict double-coordinate path
132135
auto* lrz = p.addStage<LorenzoStage<int32_t>>();
133136
// lrz->setBlockSize(32); // (7.2) block-local 1-D delta, cuSZp uses 32
134137
p.connect(lrz, quant, "codes");
135138
```
136139
140+
### Strict coordinate and reconstruction arithmetic
141+
142+
`setLinearHighPrecision(true)` evaluates `x / (2*abs_eb)` with a precomputed
143+
double reciprocal and a double multiply. Decode likewise converts the signed
144+
integer code directly to double before multiplying; converting it to float first
145+
loses low bits when `|q| > 2^24`. For arbitrary uniform steps, the stage scans
146+
the field magnitude and tightens its internal half-step by the worst-case
147+
half-ULP of the final `TInput` reconstruction. The serialized effective bound is
148+
used by decode, while the requested bound remains the external validity target.
149+
The runtime user-bound parameter is retained as double even for an f32 stage, so
150+
an f64 range calculation is not first perturbed by a float metadata conversion;
151+
strict f32 bounds are rounded downward if their representation would loosen the
152+
TOML value.
153+
154+
This accurate path is not eligible for the float-only warp-register fusion.
155+
On an H100 direct-stage NOA benchmark over 67,108,864 floats, its median forward
156+
time was 0.3794 ms versus 0.3752 ms for float coordinates (707.5 versus
157+
715.4 GB/s, 1.1% slower). The maximum error changed from 1.030x to 0.992x of
158+
the request. The benchmark is `fzgmod-profile-quantizer-linear`.
159+
160+
`setPowerOfTwoBound(true)` first resolves ABS/NOA/PREL normally and then rounds
161+
the absolute half-bound down to the nearest power of two. Thus the effective
162+
bound is never looser and is between 1x and 2x tighter. This policy follows the
163+
bound adjustment described by SLEEK: binary scaling becomes an exponent shift,
164+
avoiding general multiplication/division rounding. It is a distinct rate-
165+
distortion arm: comparisons must report the effective bound because tighter
166+
quantization can reduce compression ratio.
167+
168+
Reference: Brandon Burtchell et al., “SLEEK: An Efficient and Highly Parallel
169+
GPU Lossy Compressor,” IPDPS 2026, Section III-A,
170+
https://userweb.cs.txstate.edu/~burtscher/papers/ipdps26.pdf.
171+
172+
TOML:
173+
174+
```toml
175+
linear_mode = true
176+
linear_high_precision = true
177+
power_of_two_bound = false # true selects the separate tighter-bound arm
178+
```
179+
137180
---
138181

139182
## Error bound modes
@@ -144,8 +187,9 @@ p.connect(lrz, quant, "codes");
144187
| `NOA` | `abs(error) / value_range ≤ eb` | Scales ABS by the data range |
145188
| `REL` | `abs(error) / abs(x_orig) ≤ eb` | Ratio of error to original value |
146189

147-
**Precision.** ABS and NOA quantize, compare, and reconstruct in `TInput`, so a
148-
`double` field gets a `double` bound. REL is float32 internally (its log2/exp2
190+
**Precision.** The default ABS and NOA paths quantize, compare, and reconstruct
191+
in `TInput`; linear mode can opt into the strict double-coordinate path above.
192+
REL is float32 internally (its log2/exp2
149193
approximations are single-precision) and does not currently benefit from
150194
`TInput = double`.
151195

@@ -154,8 +198,10 @@ which for a field whose range is small next to its offset can fall below what
154198
`TInput` itself can resolve at that magnitude — the input's own round-trip through
155199
`TInput` then exceeds the bound, so no compressor could honour it. The stage compares
156200
`abs_eb` against `epsilon(TInput) * max(abs(data))` and throws rather than emitting a
157-
stream that decodes cleanly into out-of-bound values. Use a wider input type, a looser
158-
bound, or `ABS` with an explicit bound.
201+
stream that decodes cleanly into out-of-bound values. The legacy path retains a
202+
conservative spacing check; strict arbitrary-step linear mode instead reserves the
203+
exact half-ULP reconstruction budget and refuses only when that budget consumes the
204+
bound. Use a wider input type, a looser bound, or the power-of-two policy.
159205

160206
**REL mode details:**
161207
<!-- - Encodes magnitude in log2 space (PFPL), then reconstructs `x_hat` from the log bin. -->

0 commit comments

Comments
 (0)