A low-latency, GPU-resident runtime for FoundationPose — 6-DoF object pose estimation and tracking from RGB-D streams. The entire pipeline runs on the GPU with CUDA acceleration and TensorRT-optimized inference, and is exposed through a stable C ABI designed for integration from any language or framework (C/C++, Python, ROS, Rust, Go, ...). Python bindings are included.
- Two modalities: model-based (CAD mesh, OBJ/PLY) and model-free (RGB-D-mask reference views).
- Two modes:
register(one-shot global pose search) andtrack(frame-to-frame refinement at hundreds of Hz). - Multi-object: estimate poses for several objects on one shared frame concurrently.
- Zero-copy input: frames can be passed as CPU memory or CUDA device buffers (e.g. PyTorch CUDA tensors).
- Selectable inference precision: TF32 (default), strict FP32, FP16, or BF16 TensorRT engines.
- Inference micro-batching: configurable batch chunk size (
batch_size) to evaluate pose hypotheses in smaller chunks, tailoring VRAM usage to memory-constrained GPUs.
The runtime wraps a small GPU pipeline behind two integration surfaces (C ABI and Python) and keeps the two heavy stages — rendering and neural-network inference — behind swappable interfaces (a built-in CUDA rasterizer and TensorRT by default):
application ── C ABI ──┐ ┌── IRenderer ──────── built-in CUDA rasterizer
application ── Python ─┼─ estimator ┤ (optional: nvdiffrast plugin)
│ (per obj) └── IInferenceRunner ─ TensorRT
└─ persistent GPU workspace, mesh, CUDA stream
An estimator answers two calls: register (one-shot global pose search on
first sight of an object; needs a mask) and track (per-frame refinement
seeded by the previous pose; no mask). Inputs are RGB (HWC uint8), depth
(HW float32, meters), and 3×3 intrinsics; the output is a 4×4
object-to-camera SE(3) pose plus a confidence score. Frames may live in CPU
memory or CUDA device memory (auto-detected; GPU input skips the
host-to-device copy), and multiple objects can be estimated concurrently on one
shared frame.
See doc/ARCHITECTURE.md for the detailed register and
track data-flow diagrams, the C ABI / Python integration workflows
(including multi-object fp_group_*), and how to plug in a custom renderer or
inference backend through the C++ interfaces.
- NVIDIA driver ≥ 580 (the minimum driver line for CUDA 13).
- FoundationPose ONNX weights —
refiner_net.onnx,score_net.onnxfrom Hugging Facenvidia/foundationpose(orscripts/download_weights.sh). - Everything runs inside a container (
nvcr.io/nvidia/pytorch:26.05-py3, which bundles CUDA 13.2 and TensorRT 10.16); only Docker with the NVIDIA Container Toolkit is required on the host. The default build has no further dependencies — rendering uses the built-in CUDA rasterizer (seesrc/for the optional nvdiffrast backend).
Supported platforms:
| Platform | GPU | Software stack | Status |
|---|---|---|---|
| x86_64 datacenter / workstation | Compute capability ≥ 7.5 (validated on RTX PRO 6000 Blackwell; the default container build targets every arch of the base image — see src/) |
driver 580+, pytorch:26.05-py3 (CUDA 13.2, TensorRT 10.16) |
Supported |
| Jetson AGX Thor | Blackwell iGPU | JetPack 7 / L4T R38, driver 580.00, pytorch:26.05-py3 (aarch64) |
Supported — see src/ |
All containerized workflows go through run_dev.sh, which
auto-detects the host architecture and selects the right Docker Compose overlay.
Docker Compose ≥ 2.30 and the NVIDIA Container Toolkit are required on the host.
cp .env.example .env # set FP_DATA_DIR, FP_WEIGHTS_DIR, FP_UID for your machine
sed -i "s/^FP_UID.*/FP_UID=$(id -u)/" .env
sed -i "s/^FP_GID.*/FP_GID=$(id -g)/" .env./run_dev.sh build./run_dev.sh run --rm buildProduces build/libfoundation_pose_nvidia.so and build/foundation_pose_nvidia_cli.
See src/ for GPU architecture targets, renderer backends, and platform-specific notes.
scripts/download_weights.sh # ONNX weights into $FP_WEIGHTS_DIR
scripts/download_bop_ycbv.sh # BOP YCB-V dataset into $FP_DATA_DIR./run_dev.sh run --rm sample_app # C++ and Python examples (BOP mode)
./run_dev.sh run --rm sample_app python_multi # Python multi-object exampleSee example/ for all available examples, including synthetic mode
(no dataset required) and multi-object registration.
FoundationPose registration mode evaluates candidate rotation hypotheses (default: n_hypotheses = 252) across RefineNet and ScoreNet. On memory-constrained GPUs or in multi-model perception pipelines (e.g. running alongside FoundationStereo and SAM), you can tune memory usage through configuration:
By default, all hypotheses are executed in a single TensorRT batch (batch_size = 252). You can configure batch_size (via C ABI fp_config_t::batch_size, C++ Config::batch_size, or Python RuntimeConfig(batch_size=...)) to chunk TensorRT execution into smaller mini-batches (e.g. 42, 63, or 126):
- Lower VRAM footprint: RefineNet's TensorRT engine and activation buffers are allocated to fit the smaller batch size, reducing peak device memory during registration (ScoreNet uses cross-hypothesis attention and always runs at full batch).
- Full search coverage preserved: The total search coverage (
n_hypotheses) remains unchanged; RefineNet hypotheses are evaluated sequentially across chunks. Note thatbatch_sizemust dividen_hypothesesand cannot be combined withcapture_cuda_graph.
from foundation_pose_nvidia import Estimator, EstimatorOptions, RuntimeConfig
config = RuntimeConfig(
n_hypotheses=252, # total rotation search grid
batch_size=42, # evaluate in micro-batches of 42
)
with Estimator(options, config) as est:
...When compiling TensorRT engines from ONNX (refiner_net.onnx and score_net.onnx), the builder allocates temporary workspace memory (default: 8 GB, Config::tensorrt_workspace_bytes).
Lowering the workspace size on lower-memory GPUs (e.g. 8 GB–16 GB cards or embedded platforms) requires lowering batch_size accordingly. The workspace needed by TensorRT's builder scales directly with the optimization profile's batch dimension:
- In testing, building an engine with the default batch size of 252 requires at least ~6.1 GB of builder workspace.
- To successfully compile RefineNet under constrained workspace limits (e.g. 2–4 GB), pair the reduced workspace with a smaller micro-batch size (such as 42, 63, or 126). Note that ScoreNet always builds and runs at full batch.
Built-in CUDA rasterizer, FP16 precision, BOP YCB-V (4,123 targets), RTX PRO 6000 Blackwell:
| Mode | Mean latency | BOP19 AR |
|---|---|---|
| Register | 148.37 ms / target | 0.9165 |
| Tracking (video) | 2.01 ms / target (~498 Hz) | 0.9054 |
Register is one-shot per keyframe; tracking is frame-to-frame on the full video sequences (test_all).
See benchmarks/ for full latency and accuracy comparisons across TF32/FP16 precisions on both RTX PRO 6000 Blackwell and Jetson AGX Thor.
This project is licensed under Apache-2.0. Every source and
script file carries an SPDX-License-Identifier: Apache-2.0 header; see
LICENSE for the full terms.
This project downloads and installs additional third-party open source software
at build and run time (nvdiffrast, TensorRT, CUDA libraries, BOP toolkit,
Python packages, and datasets/model weights fetched by the download scripts).
Review the license terms of these projects before use — see
THIRD_PARTY_NOTICES.md for the component list and
license references. Note in particular that the optional nvdiffrast
renderer plugin is governed by the NVIDIA Source Code License and is
disabled by default: it is fetched and built only when explicitly enabled
with WITH_NVDIFFRAST=1 / -DFOUNDATION_POSE_WITH_NVDIFFRAST=ON, as its own
dynamically loaded shared library, never into the Apache-2.0 main library;
the default build contains no nvdiffrast code.
By contributing to this project, you agree that your contributions will be licensed under its Apache License, Version 2.0. The product roadmap is managed by NVIDIA.
Unless you explicitly state otherwise, any contribution intentionally submitted for inclusion in the work by you, as defined in the Apache-2.0 license, shall be dual-licensed/licensed under those terms without any additional conditions.