Skip to content

Commit 7b7e022

Browse files
committed
adds some library documentation
1 parent 7560174 commit 7b7e022

4 files changed

Lines changed: 322 additions & 0 deletions

File tree

‎CHANGELOG.md‎

Lines changed: 79 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,79 @@
1+
# Changelog
2+
3+
All notable changes to FZGPUModules are documented here.
4+
5+
Format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
6+
Version numbers follow [Semantic Versioning](https://semver.org/).
7+
8+
---
9+
10+
## [Unreleased] — 2.0.0
11+
12+
### Added
13+
14+
**Pipeline features**
15+
- Multi-source pipeline support: `InputSpec` API, `compress(std::vector<InputSpec>)`, `decompressMulti()`, `setInputSizeHint()` per source
16+
- `Pipeline::warmup(stream)` — forces PTX→SASS JIT compilation before timing-sensitive work
17+
- `Pipeline::enableBoundsCheck(bool)` — runtime toggle for buffer overwrite detection (always on in Debug builds)
18+
- `Pipeline::setCaptureMode(bool)` — CUDA Graph stream capture for steady-state compression
19+
- `Pipeline::setPoolManagedDecompOutput(bool)` — opt-in pool-owned decompression output (avoids D2D copy)
20+
- Cached inverse DAG: `buildInverseDAG()` result cached after first `decompress()` call; eliminates ~200–500 µs per-call DAG rebuild overhead
21+
- Logging system: `FZ_LOG(LEVEL, ...)` with compile-time gating; `FZ_LOG_MIN_LEVEL` CMake option (0=TRACE … 255=SILENT); `Logger::setLogCallback()` for custom sinks
22+
23+
**Memory & DAG**
24+
- Buffer coloring: non-overlapping buffers in PREALLOCATE mode are aliased to reduce peak GPU memory
25+
- Pinned concat header buffer: reduces H2D API calls from `1+N` to 1 per `compress()` call
26+
- Custom gather kernel (`launch_gather_kernel`) for D2D segment copies: replaces N individual `cudaMemcpyAsync` calls with a single kernel dispatch
27+
- `getActualOutputSize(int index)` index-based accessor on `Stage` — eliminates per-call `unordered_map` allocations in the inner execute loop
28+
- Pool auto-sizing: `computeTopoPoolSize()` + `setReleaseThreshold()` for topology-aware pool configuration
29+
- `setExternalPointer()` zero-copy path: user-owned device buffer passed directly into DAG
30+
31+
**File format**
32+
- FZM version bumped to 3 (`FZM_VERSION = 0x0300`); `FZMHeaderCore` extended to 80 bytes with `num_sources` and `source_uncompressed_sizes[4]` fields
33+
- CRC32 (IEEE 802.3) checksums on payload and header
34+
- `Pipeline::writeToFile()` / `Pipeline::decompressFromFile()` static utility
35+
36+
**Stages**
37+
- `BitshuffleStage` — GPU bit-matrix transpose
38+
- `RZEStage` — recursive zero-byte elimination with optimized kernel
39+
- `ZigzagStage<TIn, TOut>` — zigzag encode/decode
40+
- `NegabinaryStage<TIn, TOut>` — negabinary encode/decode
41+
- `DifferenceStage<T, TOut>` — first-order difference / cumulative-sum coding
42+
- `QuantizerStage<TInput, TCode>` — direct-value quantizer with ABS/REL/NOA error modes
43+
- Multi-dimensional Lorenzo (2-D and 3-D predictor kernels)
44+
45+
**Build & distribution**
46+
- `find_package(FZGPUModules REQUIRED)` support via `cmake/FZGPUModulesConfig.cmake.in`
47+
- Versioned shared library symlinks (`libfzgmod.so → .so.2 → .so.2.0.0`) via `VERSION`/`SOVERSION` target properties
48+
- Relocatable RPATH (`$ORIGIN/../lib` on Linux, `@loader_path/../lib` on macOS)
49+
- CUDA 11.2+ version floor check at CMake configure time
50+
- Little-endian host check at CMake configure time
51+
- `profiling/` directory for profiling programs (separated from `examples/`)
52+
- `scripts/new_stage.sh` scaffold script for adding new stages
53+
- CMakePresets with `asan` and `compute-san` presets for sanitizer builds
54+
- Doxygen CI GitHub Actions workflow publishing to GitHub Pages
55+
56+
**Testing**
57+
- Comprehensive test suite: 20 test binaries covering pipeline, stages, file I/O, memory strategies, buffer coloring, CUDA Graphs, bounds checking, and error handling
58+
- All tests pass under CUDA Compute Sanitizer (memcheck, initcheck, racecheck, synccheck) and host ASan+UBSan
59+
60+
### Fixed
61+
- Race condition in `CompressionDAG::execute()` for multi-source pipelines: internal per-branch streams now have a GPU-side happens-before edge into the caller stream via `cudaStreamWaitEvent`
62+
- `DifferenceStage` inverse: replaced `cub::DeviceScan` with a custom `cumsumChunkedKernel` that uses only shared memory (no device temp allocation, sanitizer-clean)
63+
- `decompressFromFile` cleanup frees use `stream=0` to avoid pool-destructor race with `cudaMemPoolDestroy`
64+
- Removed spurious `find_dependency(CCCL REQUIRED)` from installed CMake config file (would break downstream `find_package` for users without CCCL)
65+
- Removed C language from `project()` (only CUDA and CXX required)
66+
67+
### Changed
68+
- Project version set to `2.0.0` in `CMakeLists.txt`
69+
- `buildInverseDAG()` return type changed from `{inv_dag, int}` to `{inv_dag, unordered_map<Stage*, int>}` for multi-source support (breaking for internal callers)
70+
- `Stage::saveState()` / `restoreState()` contract: any stage that modifies `actual_output_sizes_` in its inverse `execute()` must override both methods
71+
72+
---
73+
74+
## [1.0.0] — 2026-04-14
75+
76+
Initial tagged release. cuSZ based compressor with modular design and experimental CUDASTF support.
77+
78+
[Unreleased]: https://github.com/szcompressor/FZGPUModules/compare/v1.1.0...HEAD
79+
[1.1.0]: https://github.com/szcompressor/FZGPUModules/releases/tag/release

‎README.md‎

Lines changed: 243 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,243 @@
1+
# FZGPUModules
2+
3+
GPU-accelerated modular lossy compression pipeline for scientific floating-point data.
4+
5+
## Overview
6+
7+
FZGPUModules is a CUDA library for building composable, high-throughput compression
8+
pipelines. Each pipeline is a directed acyclic graph (DAG) of stages — predictors,
9+
quantizers, encoders, and transforms — connected and executed entirely on the GPU
10+
with stream-ordered memory management.
11+
12+
**Key properties:**
13+
- **Modular** — mix and match stages (Lorenzo, Quantizer, RLE, RZE, Bitshuffle, …)
14+
- **High throughput** — parallel level execution, persistent scratch, CUDA Graph support
15+
- **Memory-efficient** — MINIMAL and PREALLOCATE strategies; buffer coloring to alias non-overlapping allocations
16+
17+
---
18+
19+
## Requirements
20+
21+
| Requirement | Minimum |
22+
|---|---|
23+
| CUDA Toolkit | 11.2+ (stream-ordered allocator) |
24+
| C++ Standard | C++17 |
25+
| CMake | 3.24+ |
26+
| Host byte order | Little-endian |
27+
28+
---
29+
30+
## Quick Start
31+
32+
```cpp
33+
#include "fzgpumodules.h"
34+
35+
// 1. Build a pipeline
36+
fz::Pipeline pipeline(fz::MemoryPoolConfig(input_bytes));
37+
38+
auto* lrz = pipeline.addStage<fz::LorenzoStage<float, uint16_t>>(
39+
fz::LorenzoStage<float, uint16_t>::Config{1e-4f});
40+
auto* rle = pipeline.addStage<fz::RLEStage<uint16_t>>();
41+
42+
pipeline.connect(lrz, "codes", rle);
43+
pipeline.finalize();
44+
45+
// 2. Compress
46+
void* d_compressed = nullptr;
47+
size_t compressed_size = 0;
48+
pipeline.compress(d_input, n * sizeof(float), &d_compressed, &compressed_size, stream);
49+
50+
// 3. Decompress
51+
void* d_output = nullptr;
52+
size_t output_size = 0;
53+
pipeline.decompress(d_compressed, compressed_size, &d_output, &output_size, stream);
54+
cudaStreamSynchronize(stream);
55+
// d_output is caller-owned — call cudaFree() when done.
56+
cudaFree(d_output);
57+
```
58+
59+
See `examples/` for more usage patterns including multi-branch pipelines, CUDA Graph
60+
capture, and the low-level DAG API.
61+
62+
---
63+
64+
## Caller-Allocated Output (With Size Query)
65+
66+
If you want full memory control, use the `Into` APIs and ask the pipeline for
67+
max output sizes before allocating:
68+
69+
```cpp
70+
// After finalize()
71+
size_t comp_capacity = pipeline.getMaxCompressedOutputSize();
72+
size_t decomp_capacity = pipeline.getMaxDecompressedOutputSize();
73+
74+
void* d_comp_user = nullptr;
75+
void* d_decomp_user = nullptr;
76+
cudaMalloc(&d_comp_user, comp_capacity);
77+
cudaMalloc(&d_decomp_user, decomp_capacity);
78+
79+
size_t comp_size = 0;
80+
pipeline.compressInto(d_input, input_bytes,
81+
d_comp_user, comp_capacity,
82+
&comp_size, stream);
83+
84+
size_t decomp_size = 0;
85+
pipeline.decompressInto(d_comp_user, comp_size,
86+
d_decomp_user, decomp_capacity,
87+
&decomp_size, stream);
88+
```
89+
90+
For multi-source pipelines:
91+
- Use `compressInto(const std::vector<InputSpec>&, ...)`
92+
- Use `getMaxDecompressedOutputSizes()` to size one output buffer per source
93+
- Use `decompressMultiInto(...)` to decode directly into those buffers
94+
95+
For `.fzm` files, you can query exact decompressed size(s) from header metadata:
96+
- `Pipeline::getDecompressedOutputSizeFromFile(path)`
97+
- `Pipeline::getDecompressedOutputSizesFromFile(path)`
98+
99+
See `examples/caller_allocated_output.cpp` for a minimal end-to-end example.
100+
101+
---
102+
103+
## Building from Source
104+
105+
```bash
106+
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release
107+
cmake --build build -j
108+
```
109+
110+
**CMake options:**
111+
112+
| Option | Default | Description |
113+
|---|---|---|
114+
| `BUILD_SHARED_LIBS` | `ON` | Build shared libraries |
115+
| `BUILD_EXAMPLES` | `OFF` | Build example programs |
116+
| `BUILD_PROFILING` | `OFF` | Build profiling programs |
117+
| `BUILD_TESTING` | `OFF` | Build test suite |
118+
| `FZ_LOG_MIN_LEVEL` | `2` (INFO) | 0=TRACE 1=DEBUG 2=INFO 3=WARN 255=SILENT |
119+
120+
---
121+
122+
## Installing
123+
124+
```bash
125+
cmake --install build --prefix /your/install/prefix
126+
```
127+
128+
This installs headers, shared libraries with versioned symlinks
129+
(`libfzgmod.so → libfzgmod.so.2 → libfzgmod.so.2.0.0`), and CMake
130+
package config files.
131+
132+
### Downstream CMake usage
133+
134+
```cmake
135+
find_package(FZGPUModules REQUIRED)
136+
target_link_libraries(my_target PRIVATE FZGMOD::fzgmod)
137+
```
138+
139+
All targets: `FZGMOD::fzgmod`, `FZGMOD::fzgmod_mem`, `FZGMOD::fzgmod_encoders`,
140+
`FZGMOD::fzgmod_predictors`, `FZGMOD::fzgmod_pipeline`.
141+
142+
---
143+
144+
## Available Stages
145+
146+
| Stage | Description |
147+
|---|---|
148+
| `LorenzoStage<TInput, TCode>` | 1-D/2-D/3-D Lorenzo predictor |
149+
| `QuantizerStage<TInput, TCode>` | Direct-value quantizer (ABS/REL/NOA error modes) |
150+
| `RLEStage<T>` | Run-length encoding |
151+
| `DifferenceStage<T, TOut>` | First-order difference / cumulative-sum coding |
152+
| `BitshuffleStage` | GPU bit-matrix transpose |
153+
| `RZEStage` | Recursive zero-byte elimination |
154+
| `ZigzagStage<TIn, TOut>` | Zigzag encode/decode |
155+
| `NegabinaryStage<TIn, TOut>` | Negabinary encode/decode |
156+
157+
---
158+
159+
## Memory Strategies
160+
161+
| Strategy | Description |
162+
|---|---|
163+
| `MINIMAL` | Allocate on demand, free at last consumer. Lowest peak GPU memory. |
164+
| `PREALLOCATE` | Allocate everything at `finalize()`. Required for CUDA Graph capture. Enables buffer coloring. |
165+
166+
---
167+
168+
## File I/O
169+
170+
```cpp
171+
// Write to file after compressing
172+
pipeline.writeToFile("output.fzm", stream);
173+
174+
// Decompress directly from file (no pipeline setup needed)
175+
void* d_out = nullptr;
176+
size_t out_size = 0;
177+
fz::Pipeline::decompressFromFile("output.fzm", &d_out, &out_size, stream);
178+
cudaStreamSynchronize(stream);
179+
cudaFree(d_out);
180+
```
181+
182+
FZM files embed the full stage configuration and compressed payload with CRC32
183+
checksums. See `include/fzm_format.h` for the format specification.
184+
185+
---
186+
187+
## CUDA Graph Support
188+
189+
For throughput-critical workloads, enable CUDA Graph capture to eliminate
190+
CPU-side kernel launch overhead on repeated compress calls:
191+
192+
```cpp
193+
pipeline.setCaptureMode(true); // requires PREALLOCATE strategy
194+
pipeline.finalize();
195+
pipeline.warmup(stream); // JIT-compiles all kernels once
196+
// subsequent compress() calls replay the captured graph
197+
```
198+
199+
---
200+
201+
## Key API Reference
202+
203+
| Class | Header | Description |
204+
|---|---|---|
205+
| `fz::Pipeline` | `pipeline/compressor.h` | High-level builder and executor |
206+
| `fz::Stage` | `stage/stage.h` | Base class for all compression stages |
207+
| `fz::CompressionDAG` | `pipeline/dag.h` | Low-level DAG wiring and execution |
208+
| `fz::MemoryPool` | `mem/mempool.h` | Stream-ordered CUDA memory pool |
209+
| `fz::PipelinePerfResult` | `pipeline/perf.h` | Per-stage profiling results |
210+
211+
Full API documentation is available on the [Doxygen site](https://szcompressor.github.io/FZGPUModules/).
212+
213+
---
214+
215+
## Citation
216+
217+
If you reference this work, please cite the following publication.
218+
219+
> Note: This citation corresponds to the v1.0.0 release of this codebase, not the current version.
220+
221+
- **[DRBSD-11]** FZModules: A Heterogeneous Computing Framework for Customizable Scientific Data Compression Pipelines
222+
```bibtex
223+
@inproceedings{ruiter2025fzmodules,
224+
author = {Ruiter, Skyler and Tian, Jiannan and Song, Fengguang},
225+
title = {FZModules: A Heterogeneous Computing Framework for Customizable Scientific Data Compression Pipelines},
226+
year = {2025},
227+
url = {https://doi.org/10.1145/3731599.3767376},
228+
booktitle = {Proceedings of the SC '25 Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis},
229+
pages = {332-338},
230+
series = {SC Workshops '25}
231+
}
232+
```
233+
234+
---
235+
236+
## Adding a Custom Stage
237+
238+
See `memory/how_to_add_a_stage.md` for the step-by-step guide, or use the
239+
scaffold script:
240+
241+
```bash
242+
scripts/new_stage.sh MyStageName transforms
243+
```

0 commit comments

Comments
 (0)