Skip to content

Commit 972109c

Browse files
committed
adds some more complete examples for library features and intended patterns
1 parent 26a9185 commit 972109c

15 files changed

Lines changed: 644 additions & 80 deletions

‎CHANGELOG.md‎

Lines changed: 12 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -10,6 +10,18 @@ Version numbers follow [Semantic Versioning](https://semver.org/).
1010
## [Unreleased] — 2.0.0
1111

1212
### Fixed
13+
- `src/mem/mempool.cpp` `MemoryPool::setReleaseThreshold()`: skip `config_` update when `bytes == UINT64_MAX` (the "never trim" sentinel); the previous code set `config_.input_data_size = UINT64_MAX`, causing `getConfiguredSize()` → `getPoolSize()` to overflow a float cast and return 0, which made `Pipeline::getPoolThreshold()` always return 0 post-finalize
14+
- `examples/analyze_lorenzo.cpp`, `examples/compare_lorenzo_modes.cpp`: replaced internal `#include "fused/lorenzo_quant/lorenzo_quant.h"` and `#include "mem/mempool.h"` with `#include "fzgpumodules.h"` to comply with the public-API-only rule for example code
15+
- `examples/simple_api_lorenzo_dual_branch.cpp` → `examples/lorenzo_intro.cpp`: renamed to match new catalog naming convention; CMake target renamed `simple_api_lorenzo_dual_branch` → `lorenzo_intro`; doc block updated; fixed misplaced `#include <fstream>` stranded after `using namespace fz;`; updated output header label
16+
- `examples/pfpl_pipeline.cpp` → `examples/pfpl_memory_strategies.cpp`: renamed for clarity; CMake target renamed `pfpl_example` → `pfpl_memory_strategies`; doc block reframed around the memory-strategy comparison concept; removed `getPoolThreshold()` call and table row (returns 0 post-finalize)
17+
- `examples/manual_pipeline.cpp` → `examples/pfpl_manual_vs_dag.cpp`: renamed to match the existing binary name; doc block updated to frame the "when to bypass the DAG" angle
18+
- `examples/pfpl_graph_capture.cpp`: removed `getPoolThreshold()` calls and "Pool threshold" row from comparison table (same post-finalize = 0 issue); added Build reminder to doc block
19+
- `examples/file_io_example.cpp`: fixed Usage binary path in doc block; added Build reminder; verified all four decompress paths against current API
20+
- `examples/ownership_example.cpp`: fixed Usage binary path in doc block; added Build reminder; verified all four ownership sections
21+
- `examples/debug_logging.cpp`: fixed Usage binary path in doc block; added Build reminder; removed policy-violating `#include "log.h"` (already available via `fzgpumodules.h`)
22+
- `examples/minimal_intro.cpp`: new hello-world example — LorenzoQuantStage→HuffmanStage compress+decompress+verify, synthetic data, ~80 lines
23+
- `examples/toml_config.cpp`: new TOML config example — loadConfig() from preset, saveConfig(), loadConfig() round-trip verification, Pipeline::readHeader() on .fzm file; documents the correct Pattern of Pipeline(data_bytes)+loadConfig() for PREALLOCATE presets
24+
- `examples/cusz_standalone.cpp`: new standalone stage execution example — LorenzoQuantStage+HuffmanStage driven via execute() without Pipeline; covers pool construction, onFinalize() scratch pre-allocation, postStreamSync() for outlier count, setInverse() for decompress direction
1325
- `adm_map_decoupled_u16`/`adm_map_thrust_u16`/`adm_map_decoupled_u32`/`adm_map_thrust_u32`: added `#ifndef NDEBUG` overflow sentinel — kernels call `atomicOr(d_overflow_flag, 1)` when a thread's `bit_offset` exceeds `kChunk × kMaxSignalBytes × 8`; host checks the flag after `cudaStreamSynchronize` and throws a `std::runtime_error` to catch inputs that violate the algorithm's bounded-diff assumption
1426
- `adm_map_decoupled_u16`/`adm_map_decoupled_u32`: `__shared__ excl_sum` was uninitialized for the first warp block (warp=0), causing non-deterministic writes to `d_concat_signals`; initialize to 0 at kernel entry
1527
- `tests/stages/test_adm.cpp` AD2 (`U32RoundTrip`): `make_u32_data` amplitude (±12000) exceeded the algorithm's per-thread `local_bits` capacity (64 bytes, supports max diff ≤ 4032); reduced amplitude to ±250

‎examples/CMakeLists.txt‎

Lines changed: 26 additions & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -6,16 +6,34 @@ if(NOT BUILD_EXAMPLES)
66
return()
77
endif()
88

9+
# ── Hello-world introduction: LorenzoQuant → Huffman compress+decompress+verify
10+
add_executable(minimal_intro minimal_intro.cpp)
11+
target_link_libraries(minimal_intro PRIVATE fzgmod)
12+
set_target_properties(minimal_intro PROPERTIES
13+
RUNTIME_OUTPUT_DIRECTORY "${CMAKE_BINARY_DIR}/bin/examples")
14+
15+
# ── TOML config: load preset, saveConfig, loadConfig round-trip ──────────────
16+
add_executable(toml_config toml_config.cpp)
17+
target_link_libraries(toml_config PRIVATE fzgmod)
18+
set_target_properties(toml_config PROPERTIES
19+
RUNTIME_OUTPUT_DIRECTORY "${CMAKE_BINARY_DIR}/bin/examples")
20+
21+
# ── cuSZ standalone: LorenzoQuant + Huffman without Pipeline DAG ──────────────
22+
add_executable(cusz_standalone cusz_standalone.cpp)
23+
target_link_libraries(cusz_standalone PRIVATE fzgmod)
24+
set_target_properties(cusz_standalone PROPERTIES
25+
RUNTIME_OUTPUT_DIRECTORY "${CMAKE_BINARY_DIR}/bin/examples")
26+
927
# ── Standalone Lorenzo code-distribution analyser ─────────────────────────────
1028
add_executable(analyze_lorenzo analyze_lorenzo.cpp)
1129
target_link_libraries(analyze_lorenzo PRIVATE fzgmod)
1230
set_target_properties(analyze_lorenzo PROPERTIES
1331
RUNTIME_OUTPUT_DIRECTORY "${CMAKE_BINARY_DIR}/bin/examples")
1432

1533
# ── PREALLOCATE vs MINIMAL memory strategy comparison ─────────────────────────
16-
add_executable(pfpl_example pfpl_pipeline.cpp)
17-
target_link_libraries(pfpl_example PRIVATE fzgmod)
18-
set_target_properties(pfpl_example PROPERTIES
34+
add_executable(pfpl_memory_strategies pfpl_memory_strategies.cpp)
35+
target_link_libraries(pfpl_memory_strategies PRIVATE fzgmod)
36+
set_target_properties(pfpl_memory_strategies PROPERTIES
1937
RUNTIME_OUTPUT_DIRECTORY "${CMAKE_BINARY_DIR}/bin/examples")
2038

2139
# ── MINIMAL / PREALLOCATE / CUDA Graph capture three-way comparison ───────────
@@ -24,14 +42,14 @@ target_link_libraries(pfpl_graph_capture PRIVATE fzgmod)
2442
set_target_properties(pfpl_graph_capture PROPERTIES
2543
RUNTIME_OUTPUT_DIRECTORY "${CMAKE_BINARY_DIR}/bin/examples")
2644

27-
# ── Simple API: Lorenzo dual-branch pipeline on a real input file ─────────────
28-
add_executable(simple_api_lorenzo_dual_branch simple_api_lorenzo_dual_branch.cpp)
29-
target_link_libraries(simple_api_lorenzo_dual_branch PRIVATE fzgmod)
30-
set_target_properties(simple_api_lorenzo_dual_branch PROPERTIES
45+
# ── Getting-started: LorenzoQuant → Bitshuffle → RZE ─────────────────────────
46+
add_executable(lorenzo_intro lorenzo_intro.cpp)
47+
target_link_libraries(lorenzo_intro PRIVATE fzgmod)
48+
set_target_properties(lorenzo_intro PROPERTIES
3149
RUNTIME_OUTPUT_DIRECTORY "${CMAKE_BINARY_DIR}/bin/examples")
3250

3351
# ── Manual stage execution vs Pipeline DAG throughput comparison ──────────────
34-
add_executable(pfpl_manual_vs_dag manual_pipeline.cpp)
52+
add_executable(pfpl_manual_vs_dag pfpl_manual_vs_dag.cpp)
3553
target_link_libraries(pfpl_manual_vs_dag PRIVATE fzgmod)
3654
set_target_properties(pfpl_manual_vs_dag PROPERTIES
3755
RUNTIME_OUTPUT_DIRECTORY "${CMAKE_BINARY_DIR}/bin/examples")

‎examples/analyze_lorenzo.cpp‎

Lines changed: 1 addition & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -24,10 +24,7 @@
2424
* Build: appears automatically in the CMake build as target "analyze_lorenzo"
2525
*/
2626

27-
#include "fused/lorenzo_quant/lorenzo_quant.h"
28-
#include "mem/mempool.h"
29-
30-
#include <cuda_runtime.h>
27+
#include "fzgpumodules.h"
3128

3229
#include <algorithm>
3330
#include <cmath>

‎examples/compare_lorenzo_modes.cpp‎

Lines changed: 1 addition & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -35,10 +35,8 @@
3535
* Binary: build/release/bin/examples/compare_lorenzo_modes
3636
*/
3737

38-
#include "fused/lorenzo_quant/lorenzo_quant.h"
39-
#include "mem/mempool.h"
38+
#include "fzgpumodules.h"
4039

41-
#include <cuda_runtime.h>
4240
#include <algorithm>
4341
#include <cmath>
4442
#include <cstdint>

‎examples/cusz_standalone.cpp‎

Lines changed: 197 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,197 @@
1+
/**
2+
* examples/cusz_standalone.cpp
3+
*
4+
* LorenzoQuantStage → HuffmanStage driven manually without a Pipeline.
5+
*
6+
* Use this pattern when you need precise control over execution order,
7+
* intermediate buffer placement, or per-stage stream assignment — for
8+
* instance, to interleave other GPU work between compression stages,
9+
* or to run stages asynchronously on multiple streams.
10+
*
11+
* Demonstrates:
12+
* - Constructing a caller-managed MemoryPool for stage scratch
13+
* - HuffmanStage::onFinalize() to pre-allocate persistent scratch buffers
14+
* - LorenzoQuantStage::execute() — 1 input, 4 outputs
15+
* (codes, outlier_errors, outlier_indices, outlier_count)
16+
* - LorenzoQuantStage::postStreamSync() to read the actual outlier count
17+
* - HuffmanStage::execute() in compress and decompress direction
18+
* - Stage::setInverse(true) to switch to the decompression pass
19+
* - LorenzoQuantStage::execute() in decompress mode — 4 inputs, 1 output
20+
* - Verifying the full compress → decompress round-trip
21+
*
22+
* Contrast with the Pipeline DAG path:
23+
* - Pipeline manages the pool, allocates intermediate buffers, and calls
24+
* postStreamSync() automatically — you write connect() instead of buffer ptrs.
25+
* - The standalone path gives you raw pointers but requires manual management
26+
* of everything the Pipeline normally hides.
27+
*
28+
* No external data files required — uses synthetic float data.
29+
*
30+
* Usage:
31+
* ./build/release/bin/examples/cusz_standalone
32+
*
33+
* Build:
34+
* cmake --preset release -DBUILD_EXAMPLES=ON && cmake --build build/release -j$(nproc)
35+
* Binary: build/release/bin/examples/cusz_standalone
36+
*/
37+
38+
#include "fzgpumodules.h"
39+
40+
#include <algorithm>
41+
#include <cmath>
42+
#include <cstdio>
43+
#include <vector>
44+
45+
using namespace fz;
46+
47+
int main() {
48+
// ── 1. Input data ─────────────────────────────────────────────────────────
49+
static constexpr size_t N = 1 << 18; // 256 K floats = 1 MB
50+
static constexpr float EB = 1e-4f;
51+
static constexpr float OUTLIER_CAPACITY = 0.10f;
52+
static constexpr uint16_t QUANT_RADIUS = 512;
53+
// Zigzag maps [-511, 511] → [0, 1022]; bklen=1024 covers the full range.
54+
static constexpr uint16_t BKLEN = 2 * QUANT_RADIUS;
55+
56+
std::vector<float> h_input(N);
57+
for (size_t i = 0; i < N; ++i) {
58+
const float t = static_cast<float>(i) / static_cast<float>(N);
59+
h_input[i] = std::sin(2.0f * 3.14159265f * t)
60+
+ 0.5f * std::cos(6.0f * 3.14159265f * t);
61+
}
62+
const size_t input_bytes = N * sizeof(float);
63+
const size_t codes_bytes = N * sizeof(uint16_t);
64+
const size_t max_outliers = static_cast<size_t>(N * OUTLIER_CAPACITY) + 1;
65+
const size_t huf_out_cap = codes_bytes * 2 + 4096; // generous upper bound
66+
67+
float* d_input = nullptr;
68+
cudaMalloc(&d_input, input_bytes);
69+
cudaMemcpy(d_input, h_input.data(), input_bytes, cudaMemcpyHostToDevice);
70+
71+
// ── 2. Stage setup ────────────────────────────────────────────────────────
72+
LorenzoQuantStage<float, uint16_t> lq;
73+
lq.setErrorBound(EB);
74+
lq.setErrorBoundMode(ErrorBoundMode::ABS);
75+
lq.setQuantRadius(QUANT_RADIUS);
76+
lq.setOutlierCapacity(OUTLIER_CAPACITY);
77+
lq.setZigzagCodes(true);
78+
79+
HuffmanStage<uint16_t> huf;
80+
huf.setBklen(BKLEN);
81+
82+
// ── 3. Pool and persistent scratch ───────────────────────────────────────
83+
//
84+
// The pool serves two roles:
85+
// a) Huffman persistent scratch: pre-allocated here via onFinalize()
86+
// so that no allocation happens inside the timed compress() loop.
87+
// b) LorenzoQuant transient scratch: tiny workspace for the NOA/REL
88+
// value-range scan (not needed here with ABS mode, but the pool
89+
// pointer is required by execute() regardless).
90+
//
91+
// Size the pool to cover Huffman's full device footprint plus padding.
92+
MemoryPool pool(MemoryPoolConfig(input_bytes, 4.0f));
93+
huf.onFinalize(codes_bytes, &pool); // pre-allocates PHF scratch buffers
94+
95+
// ── 4. Intermediate device buffers ────────────────────────────────────────
96+
//
97+
// Caller owns all of these — cudaFree every one on exit.
98+
// These are distinct from pool memory (which is internal scratch).
99+
void* d_codes = nullptr;
100+
void* d_outlier_errors = nullptr;
101+
void* d_outlier_indices = nullptr;
102+
void* d_outlier_count = nullptr;
103+
void* d_huf_output = nullptr;
104+
void* d_reconstructed = nullptr;
105+
106+
cudaMalloc(&d_codes, codes_bytes);
107+
cudaMalloc(&d_outlier_errors, max_outliers * sizeof(float));
108+
cudaMalloc(&d_outlier_indices, max_outliers * sizeof(uint32_t));
109+
cudaMalloc(&d_outlier_count, sizeof(uint32_t));
110+
cudaMalloc(&d_huf_output, huf_out_cap);
111+
cudaMalloc(&d_reconstructed, input_bytes);
112+
cudaDeviceSynchronize();
113+
114+
// ── 5. Compress: LorenzoQuant → Huffman ───────────────────────────────────
115+
//
116+
// LorenzoQuant compress: 1 input → 4 outputs.
117+
// outputs[3] = outlier_count (device uint32_t) — initialized to 0 by execute().
118+
lq.execute(0, &pool,
119+
{d_input},
120+
{d_codes, d_outlier_errors, d_outlier_indices, d_outlier_count},
121+
{input_bytes});
122+
cudaDeviceSynchronize();
123+
124+
// postStreamSync() reads the actual outlier count back to host and trims
125+
// actual_output_sizes_ to the real (not max-capacity) sizes.
126+
// Must be called after the stream is fully synchronized.
127+
lq.postStreamSync(0);
128+
129+
const size_t actual_outlier_errors_bytes = lq.getActualOutputSize(1);
130+
const size_t actual_outlier_indices_bytes = lq.getActualOutputSize(2);
131+
const size_t n_outliers = actual_outlier_errors_bytes / sizeof(float);
132+
std::printf("LorenzoQuant compress: %zu elements, %zu outliers (%.2f%%)\n",
133+
N, n_outliers, 100.0 * n_outliers / N);
134+
135+
// Huffman compress: 1 input (codes) → 1 output (compressed bitstream).
136+
huf.execute(0, &pool,
137+
{d_codes},
138+
{d_huf_output},
139+
{codes_bytes});
140+
cudaDeviceSynchronize();
141+
142+
const size_t compressed_bytes = huf.getActualOutputSize(0);
143+
std::printf("Huffman compress: %zu bytes → %zu bytes (%.2fx for codes)\n",
144+
codes_bytes, compressed_bytes,
145+
static_cast<double>(codes_bytes) / compressed_bytes);
146+
std::printf("Full ratio (input/compressed_codes): %.2fx\n",
147+
static_cast<double>(input_bytes) / compressed_bytes);
148+
149+
// ── 6. Decompress: Huffman inverse → LorenzoQuant inverse ─────────────────
150+
//
151+
// Both stages reuse the same objects with setInverse(true).
152+
// The Huffman header embedded in d_huf_output tells the decoder the
153+
// original length — no external metadata needed for the codes buffer.
154+
155+
// Huffman decompress: compressed bitstream → codes.
156+
huf.setInverse(true);
157+
huf.execute(0, &pool,
158+
{d_huf_output},
159+
{d_codes}, // reuse d_codes as output
160+
{compressed_bytes});
161+
cudaDeviceSynchronize();
162+
163+
// LorenzoQuant decompress: 4 inputs → 1 output (reconstructed floats).
164+
// sizes[1] = outlier_errors_bytes is used to derive max_outliers inside
165+
// the kernel, so pass the actual (not max-capacity) sizes.
166+
lq.setInverse(true);
167+
lq.execute(0, &pool,
168+
{d_codes, d_outlier_errors, d_outlier_indices, d_outlier_count},
169+
{d_reconstructed},
170+
{codes_bytes,
171+
actual_outlier_errors_bytes,
172+
actual_outlier_indices_bytes,
173+
sizeof(uint32_t)});
174+
cudaDeviceSynchronize();
175+
176+
// ── 7. Verify ─────────────────────────────────────────────────────────────
177+
std::vector<float> h_out(N);
178+
cudaMemcpy(h_out.data(), d_reconstructed, input_bytes, cudaMemcpyDeviceToHost);
179+
180+
float max_err = 0.0f;
181+
for (size_t i = 0; i < N; ++i)
182+
max_err = std::max(max_err, std::abs(h_out[i] - h_input[i]));
183+
184+
std::printf("max_abs_error: %.2e (eb = %.2e)\n",
185+
max_err, static_cast<double>(EB));
186+
187+
// ── 8. Cleanup ────────────────────────────────────────────────────────────
188+
cudaFree(d_input);
189+
cudaFree(d_codes);
190+
cudaFree(d_outlier_errors);
191+
cudaFree(d_outlier_indices);
192+
cudaFree(d_outlier_count);
193+
cudaFree(d_huf_output);
194+
cudaFree(d_reconstructed);
195+
// pool is stack-allocated — its destructor frees Huffman's persistent scratch.
196+
return 0;
197+
}

‎examples/debug_logging.cpp‎

Lines changed: 17 additions & 9 deletions
Original file line numberDiff line numberDiff line change
@@ -1,36 +1,44 @@
11
/**
2-
* debug_logging.cpp — debugging and logging tools for pipeline development.
2+
* examples/debug_logging.cpp
33
*
4-
* Covers the four main tools for understanding what the pipeline is doing:
4+
* Debugging and logging tools for pipeline development.
5+
*
6+
* Covers the five main tools for understanding what the pipeline is doing:
57
*
68
* 1. Logger — route library log output to stderr or a custom callback.
79
* Log levels: TRACE (0) < DEBUG (1) < INFO (2) < WARN (3) < SILENT.
810
* Default: INFO at runtime, compile-time floor set by FZ_LOG_MIN_LEVEL.
911
*
10-
* 2. printPipeline() / printDAG() / printBufferLifetimes()
12+
* 2. Custom log callback — filter, tag, or redirect log lines to any sink
13+
* (test buffer, file, GUI panel, etc.).
14+
*
15+
* 3. printPipeline() / getDAG()->printDAG() / getDAG()->printBufferLifetimes()
1116
* Explicit diagnostic dumps — always produce output regardless of
1217
* log level. Route through the Logger callback if one is set,
1318
* otherwise write to stdout.
1419
*
15-
* 3. enableProfiling() + getLastPerfResult()
20+
* 4. enableProfiling() + getLastPerfResult()
1621
* CUDA-event per-stage timing. Use to find which stage dominates,
1722
* compare pipeline variants, or profile before optimising.
1823
*
19-
* 4. enableBoundsCheck()
24+
* 5. enableBoundsCheck()
2025
* Runtime check that no stage writes beyond its allocated buffer.
2126
* Enabled automatically in debug builds; opt-in in release.
2227
*
23-
* No external data files required.
28+
* No external data files required — uses synthetic float data.
2429
*
2530
* Usage:
26-
* ./build/bin/debug_logging
31+
* ./build/bin/examples/debug_logging
2732
*
2833
* To see TRACE + DEBUG logs, rebuild with:
29-
* cmake -DFZ_LOG_MIN_LEVEL=0 --preset release && cmake --build --preset release
34+
* cmake -DFZ_LOG_MIN_LEVEL=0 --preset release && cmake --build build/release -j$(nproc)
35+
*
36+
* Build:
37+
* cmake --preset release -DBUILD_EXAMPLES=ON && cmake --build build -j$(nproc)
38+
* Binary: build/bin/examples/debug_logging
3039
*/
3140

3241
#include "fzgpumodules.h"
33-
#include "log.h"
3442

3543
#include <algorithm>
3644
#include <cmath>

0 commit comments

Comments
 (0)