You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit f5b2f10
Browse filesBrowse the repository at this point in the historyBrowse files
Add strict high-precision linear quantizer arithmetic and power-of-two bound policy
linear_high_precision keeps both coordinates and the user bound in double,
rounds strict f32 metadata downward, and reserves reconstruction rounding so a
tight requested bound is no longer violated. Optional power_of_two_bound rounds
the resolved ABS/NOA/PREL half-bound down to a power of two. Both flags persist
in the existing 72-byte header, round-trip through TOML, and stay opt-in. Adds
the fzgmod-profile-quantizer-linear four-way comparison benchmark.
Copy file name to clipboardExpand all lines: CHANGELOG.md
+2Lines changed: 2 additions & 0 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -10,6 +10,7 @@ Version numbers follow [Semantic Versioning](https://semver.org/).
10
10
## [Unreleased] — 2.0.0
11
11
12
12
### Added
13
+
-**Strict linear-Quantizer arithmetic and an optional SLEEK-style power-of-two bound policy.**`linear_high_precision` uses a precomputed double reciprocal on encode, direct integer-to-double reconstruction on decode, and an internal half-ULP reserve for arbitrary steps; `power_of_two_bound` rounds the resolved ABS/NOA/PREL half-bound downward (1x–2x tighter). Both flags persist in the existing 72-byte header by using previously-zero padding, round-trip through TOML, and remain opt-in. Added `fzgmod-profile-quantizer-linear` for the four-way float/double/power2 comparison.
13
14
- **Explicit-ownership execution API on `Pipeline` — ownership is now expressed by the return type instead of by `void**` vs `void*` and by mutable pipeline state.** New backend-neutral value types in `include/pipeline/device_buffer.h` (`DeviceSpan`, `ConstDeviceSpan`, non-owning `BorrowedDeviceBuffer`, move-only `OwnedDeviceBuffer` that records its device and frees through the backend facade, never a hard-coded `cudaFree`), plus span wrappers `compress()`, `compressInto()`, `decompressBorrowed()`, `decompressOwned()`, `decompressInto()`, and `decompressIntoAsync()`. They delegate to the existing pointer overloads, so behavior is unchanged and nothing is deprecated; `decompressBorrowed()`/`decompressOwned()` deliberately ignore `setPoolManagedDecompOutput()` and restore the flag. Tests: `tests/pipeline/test_device_buffer.cpp` (8 cases: move-only/non-owning traits, borrow round-trip, byte-identical `compressInto`, owned-buffer reclamation, ownership independent of the pipeline flag, capacity failure, move/release). 53/53 ctest. Step 1-2 of `memory/public_api_evolution.md`.
14
15
- Added `fzgmod-profile-huffman-throughput`, a direct-stage HostCoordinated/DeviceResident latency and throughput benchmark with cold/warm codebook cases, CUDA-event and host timing, multiple sizes/distributions, round-trip checks, deterministic archive comparison, and CSV output.
15
16
- Added `HuffmanExecutionMode::DeviceResident`: cuSZ-compatible canonical tree/forward/reverse-book construction directly from device histograms, device-side partition scan and PHF assembly for PerBlock/Adaptive/Fixed books, safe range validation, device-side header parsing on decode, terminal fixed-book CUDA Graph support, and automatic exact-size readback when Huffman feeds a downstream stage.
@@ -21,6 +22,7 @@ Version numbers follow [Semantic Versioning](https://semver.org/).
21
22
- Removed Huffman's experimental `Fine` encode mode, its public API and diagnostics, TOML/card option, documentation, profiling selection, and mode-specific tests; `HuffmanStage` now exposes only the supported cuSZ coarse-grained encoder.
22
23
23
24
### Fixed
25
+
-**Tight linear quantization could still violate the bound after fixing integer overflow.** Float reciprocal multiplication selected adjacent bins, final f32 reconstruction consumed additional error, the inverse first cast signed codes to float—dropping low bits above 2^24—and the runtime user bound was narrowed to float before an f64 range multiplication. The strict path now keeps both coordinate directions and the user-bound parameter in double, rounds strict f32 metadata downward rather than loosening it, and reserves reconstruction rounding. A captured CESM FLNTC `rel_range=1e-7` case changed from a 1.744x miss to bit-exact reconstruction; an 18-cell CESM/HURR/NYX smoke matrix passes every requested bound.
24
26
-**Linear quantization silently wrapped bins outside signed `TCode` range.** Tight NOA bounds on offset-valued fields can require indices much wider than int32 (S3D `N2` at `1e-6` needs roughly 3.3e10). The linear kernel now uses a guarded int64 rounding intermediate, records any signed-code overflow in device scratch, and `postStreamSync()` refuses the run with remediation guidance instead of producing a corrupt archive. Regression test: `QuantizerLinear.RefusesCodeOverflow`; rationale: `CN-QUANT-3`.
25
27
- Docker image builds now retain `THIRD_PARTY.md` in the build context so the CMake install step can package the required third-party notices.
26
28
- Installed packages now include the project `LICENSE` and `THIRD_PARTY.md`; synchronized all 30-stage documentation/catalog surfaces, completed newer-stage attribution, and corrected the cuSZ-Hi/SZx/ROIBIN-SZ license records.
Copy file name to clipboardExpand all lines: docs/cards/fzgm-quantizer.json
+12Lines changed: 12 additions & 0 deletions
Original file line number
Diff line number
Diff line change
@@ -60,6 +60,18 @@
60
60
"type": "bool",
61
61
"default": false,
62
62
"description": "ABS/NOA: linear/no-outlier mode (cuSZp-style). Emits raw signed codes q=round(x/2·eb), no radius clamp, no outlier ports. Pair with LorenzoStage(setBlockSize) → AdaptiveBitpackStage."
63
+
},
64
+
{
65
+
"name": "linearHighPrecision",
66
+
"type": "bool",
67
+
"default": false,
68
+
"description": "Linear mode only: double coordinate and reconstruction arithmetic plus an internal rounding reserve for a strict TInput error bound."
69
+
},
70
+
{
71
+
"name": "powerOfTwoBound",
72
+
"type": "bool",
73
+
"default": false,
74
+
"description": "ABS/NOA/PREL: round the resolved absolute half-bound downward to a power of two (1x to 2x tighter than requested), following SLEEK's bound policy."
0 commit comments