This document covers clean clone setup, submodules, build profiles, backend builds, and common build failures.
Preferred clone:
git clone --recursive https://github.com/THU-MIG/edge-dit.cpp
cd edge-dit.cppIf the checkout already exists:
bash scripts/bootstrap.shWhen updating an existing checkout:
git pull --recurse-submodules
git submodule update --init --recursivescripts/bootstrap.sh verifies that the current directory is a Git checkout,
runs:
git submodule update --init --recursiveand checks that third_party/ggml/CMakeLists.txt exists.
GitHub auto-generated Source ZIP archives usually do not include submodule
contents. Use git clone --recursive, scripts/bootstrap.sh, or a complete
project release source package.
Common tools:
- CMake 3.20 or newer.
- C and C++ compilers with C++17 support.
- Git, including submodule support.
Optional backend-specific tools:
- CUDA Toolkit and
nvccfor CUDA (CUDA 12.x or 13.x). - NCCL, MPI, cuDNN, and cudnn-frontend for the official CUDA performance profile.
- Apple build tools and frameworks for Metal on macOS.
- Vulkan SDK,
glslc, Vulkan headers, and SPIRV-Headers for Vulkan. - Python 3 for Python bindings and release tooling.
The build scripts honor standard overrides:
CMAKE_BIN
CC
CXX
CUDACXX
CUDA_HOME
CUDA_PATH
NCCL_ROOT
CUDNN_ROOT
MPI_HOME
BUILD_DIR
CLEAN
CUDA builds use ED_BUILD_PROFILE.
performance is the default and official performance configuration.
It enables the CUDA backend and the single-GPU performance optimizations:
- cuDNN SDPA
- CUDA Norm
- CUDA RoPE
- CUDA Modulation
- benchmark phase markers (
ED_BENCHMARK_MARKERS, for component-level timing)
Multi-GPU dependencies -- NCCL, MPI, and parallel runtime
(ED_ENABLE_NCCL / ED_ENABLE_MPI / ED_ENABLE_PARALLEL) -- default to
OFF, so a single-GPU build needs neither NCCL nor MPI installed. Enable them
explicitly for multi-GPU / distributed runs:
ED_ENABLE_NCCL=ON NCCL_ROOT=/path/to/nccl \
ED_ENABLE_MPI=ON MPI_HOME=/path/to/mpi \
ED_ENABLE_PARALLEL=ON \
ED_BUILD_PROFILE=performance \
bash scripts/build_cuda.shIf an explicitly enabled dependency is missing, the script fails before CMake configure. It does not silently downgrade to a slower configuration. The default single-GPU build needs no extra flags:
bash scripts/build_cuda.shperformance remains the default:
bash scripts/build_cuda.shFull performance profile validation is the v0.1.0 release gate.
Do not treat minimal results as official performance data.
minimal is a reduced external dependency configuration for build and
functional validation:
ED_BUILD_PROFILE=minimal bash scripts/build_cuda.shIt disables NCCL, MPI, and cuDNN SDPA by default. It keeps the CUDA backend and the in-repository CUDA helper paths needed by the current patched ggml CUDA integration. It is not valid for official benchmark results.
Environment variables set by the user override profile defaults. For example:
ED_BUILD_PROFILE=performance \
ED_ENABLE_CUDNN_SDPA=OFF \
bash scripts/build_cuda.shThe CUDA build script is intended to be beginner-friendly. By default it tries to install or fetch dependencies that can be handled safely in user space:
bash scripts/build_cuda.shFor performance builds this means:
ED_INSTALL_CUDNN=ONby default, using user-level NVIDIA Python wheels when cuDNN is missing andCUDNN_ROOTis not set.ED_INSTALL_CUDNN_FRONTEND=ONby default, which setsED_FETCH_CUDNN_FRONTEND=ONwhen the vendored cudnn-frontend source is absent.
Disable automatic user-space dependency installation with:
ED_AUTO_INSTALL_DEPS=OFF bash scripts/build_cuda.shCUDA Toolkit, NCCL, and MPI remain system-level dependencies. The script does
not silently install drivers, compilers, MPI implementations, or NCCL system
packages. Install them with your environment's package manager or set
CUDA_HOME, NCCL_ROOT, and MPI_HOME.
Before CMake configure, scripts/build_cuda.sh prints a compact summary with:
- profile
- source and build directories
- CMake and compiler paths / versions
- CUDA Toolkit and
nvcc - CUDA architectures
- NCCL, MPI, and cuDNN status
- cuDNN SDPA, CUDA Norm, CUDA RoPE, CUDA Modulation
- CFG and sequence parallel status
- build type and library mode
After configure, CMake writes:
<build-dir>/build-config.txt
The file records resolved feature switches, dependency versions when available, CUDA architectures, compiler versions, build type, Git commit, and ggml submodule commit. It intentionally avoids unnecessary personal paths.
bash scripts/build_cpu.shCPU is mainly for build validation, smoke tests, fallback operators, CPU offload, and selected low-speed inference.
Output defaults to:
build-cpu/
Quick validation:
ED_BUILD_PROFILE=minimal bash scripts/build_cuda.shPerformance validation:
CUDA_HOME=/path/to/cuda \
NCCL_ROOT=/path/to/nccl \
CUDNN_ROOT=/path/to/cudnn \
MPI_HOME=/path/to/mpi \
ED_BUILD_PROFILE=performance \
BUILD_DIR=build-cuda-performance \
bash scripts/build_cuda.shOutput defaults to:
build-cuda/
Metal is macOS-only and experimental:
bash scripts/build_metal.shOutput defaults to:
build-metal/
Vulkan is experimental. The helper script looks for glslc, Vulkan headers,
and SPIRV-Headers:
bash scripts/build_vulkan.shUseful overrides:
VULKAN_SDK=/path/to/vulkan-sdk bash scripts/build_vulkan.sh
VK_EXTRA_HEADERS=/path/to/Vulkan-Headers/include \
SPIRV_HEADERS=/path/to/SPIRV-Headers/include \
bash scripts/build_vulkan.shOutput defaults to:
build-vulkan/
Each build script accepts:
CLEAN=1 bash scripts/build_cpu.shCLEAN=1 only removes the selected build directory.
Symptom:
Missing ggml submodule at third_party/ggml
Fix:
bash scripts/bootstrap.shSymptom:
ED_ENABLE_NCCL=ON requires NCCL headers and library
Fix:
NCCL_ROOT=/path/to/nccl ED_BUILD_PROFILE=performance bash scripts/build_cuda.shInstall an MPI implementation or set:
MPI_HOME=/path/to/mpiInstall cuDNN and set:
CUDNN_ROOT=/path/to/cudnnIf the vendored cudnn-frontend source is absent, either initialize submodules or explicitly allow configure-time fetching:
ED_FETCH_CUDNN_FRONTEND=ONFrom the repository root:
PYTHONPATH=bindings/python/src python3 -m pytest bindings/python/tests