You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 6f82fbf
Browse filesBrowse the repository at this point in the historyBrowse files
Wires the GPU serving path behind -DSQLITE_PREDICT_ONNX_GPU
(make loadable-onnx-gpu):
- CUDA and TensorRT execution providers via proper provider-options
(Create/Update/Release), not the previous NULL stub. TensorRT honors
trt_fp16_enable when precision=fp16. Appending fails loud if the EP is
not in the onnxruntime build — never a silent drop to CPU.
- fp16/int8 precision allowed only in the GPU build, and only with a
cuda/tensorrt device (pairing fp16 with cpu/coreml is rejected rather
than silently computing fp32).
- The provider-options symbols are in every onnxruntime C API, so the GPU
build compiles and links against the CPU onnxruntime. CI now compile-
checks it (make loadable-onnx-gpu); real GPU execution is validated on a
dedicated GPU job (needs onnxruntime-gpu + a GPU runner).
Verified locally against the CPU onnxruntime: the GPU build compiles and
links clean, device=cuda/tensorrt fail loud with RUNTIME_UNAVAILABLE (no
crash), fp16+cpu is rejected, and the regular onnx suite stays 124 green.
The receipt already records device+precision, so GPU results are honestly
distinguishable from the deterministic CPU path. README/ARCHITECTURE/
CHANGELOG updated.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: mstrathman <matthew.strathman@gmail.com>
0 commit comments