Skip to content

Latest commit

 

History

History
352 lines (247 loc) · 7.23 KB

File metadata and controls

352 lines (247 loc) · 7.23 KB

Build and Installation

← Back to README

This document covers clean clone setup, submodules, build profiles, backend builds, and common build failures.

Clean Clone

Preferred clone:

git clone --recursive https://github.com/THU-MIG/edge-dit.cpp
cd edge-dit.cpp

If the checkout already exists:

bash scripts/bootstrap.sh

When updating an existing checkout:

git pull --recurse-submodules
git submodule update --init --recursive

scripts/bootstrap.sh verifies that the current directory is a Git checkout, runs:

git submodule update --init --recursive

and checks that third_party/ggml/CMakeLists.txt exists.

GitHub auto-generated Source ZIP archives usually do not include submodule contents. Use git clone --recursive, scripts/bootstrap.sh, or a complete project release source package.

Prerequisites

Common tools:

  • CMake 3.20 or newer.
  • C and C++ compilers with C++17 support.
  • Git, including submodule support.

Optional backend-specific tools:

  • CUDA Toolkit and nvcc for CUDA (CUDA 12.x or 13.x).
  • NCCL, MPI, cuDNN, and cudnn-frontend for the official CUDA performance profile.
  • Apple build tools and frameworks for Metal on macOS.
  • Vulkan SDK, glslc, Vulkan headers, and SPIRV-Headers for Vulkan.
  • Python 3 for Python bindings and release tooling.

The build scripts honor standard overrides:

CMAKE_BIN
CC
CXX
CUDACXX
CUDA_HOME
CUDA_PATH
NCCL_ROOT
CUDNN_ROOT
MPI_HOME
BUILD_DIR
CLEAN

CUDA Build Profiles

CUDA builds use ED_BUILD_PROFILE.

performance

performance is the default and official performance configuration.

It enables the CUDA backend and the single-GPU performance optimizations:

  • cuDNN SDPA
  • CUDA Norm
  • CUDA RoPE
  • CUDA Modulation
  • benchmark phase markers (ED_BENCHMARK_MARKERS, for component-level timing)

Multi-GPU dependencies -- NCCL, MPI, and parallel runtime (ED_ENABLE_NCCL / ED_ENABLE_MPI / ED_ENABLE_PARALLEL) -- default to OFF, so a single-GPU build needs neither NCCL nor MPI installed. Enable them explicitly for multi-GPU / distributed runs:

ED_ENABLE_NCCL=ON NCCL_ROOT=/path/to/nccl \
ED_ENABLE_MPI=ON MPI_HOME=/path/to/mpi \
ED_ENABLE_PARALLEL=ON \
ED_BUILD_PROFILE=performance \
bash scripts/build_cuda.sh

If an explicitly enabled dependency is missing, the script fails before CMake configure. It does not silently downgrade to a slower configuration. The default single-GPU build needs no extra flags:

bash scripts/build_cuda.sh

performance remains the default:

bash scripts/build_cuda.sh

Full performance profile validation is the v0.1.0 release gate. Do not treat minimal results as official performance data.

minimal

minimal is a reduced external dependency configuration for build and functional validation:

ED_BUILD_PROFILE=minimal bash scripts/build_cuda.sh

It disables NCCL, MPI, and cuDNN SDPA by default. It keeps the CUDA backend and the in-repository CUDA helper paths needed by the current patched ggml CUDA integration. It is not valid for official benchmark results.

User overrides

Environment variables set by the user override profile defaults. For example:

ED_BUILD_PROFILE=performance \
ED_ENABLE_CUDNN_SDPA=OFF \
bash scripts/build_cuda.sh

The CUDA build script is intended to be beginner-friendly. By default it tries to install or fetch dependencies that can be handled safely in user space:

bash scripts/build_cuda.sh

For performance builds this means:

  • ED_INSTALL_CUDNN=ON by default, using user-level NVIDIA Python wheels when cuDNN is missing and CUDNN_ROOT is not set.
  • ED_INSTALL_CUDNN_FRONTEND=ON by default, which sets ED_FETCH_CUDNN_FRONTEND=ON when the vendored cudnn-frontend source is absent.

Disable automatic user-space dependency installation with:

ED_AUTO_INSTALL_DEPS=OFF bash scripts/build_cuda.sh

CUDA Toolkit, NCCL, and MPI remain system-level dependencies. The script does not silently install drivers, compilers, MPI implementations, or NCCL system packages. Install them with your environment's package manager or set CUDA_HOME, NCCL_ROOT, and MPI_HOME.

CUDA Configuration Summary

Before CMake configure, scripts/build_cuda.sh prints a compact summary with:

  • profile
  • source and build directories
  • CMake and compiler paths / versions
  • CUDA Toolkit and nvcc
  • CUDA architectures
  • NCCL, MPI, and cuDNN status
  • cuDNN SDPA, CUDA Norm, CUDA RoPE, CUDA Modulation
  • CFG and sequence parallel status
  • build type and library mode

After configure, CMake writes:

<build-dir>/build-config.txt

The file records resolved feature switches, dependency versions when available, CUDA architectures, compiler versions, build type, Git commit, and ggml submodule commit. It intentionally avoids unnecessary personal paths.

CPU Build

bash scripts/build_cpu.sh

CPU is mainly for build validation, smoke tests, fallback operators, CPU offload, and selected low-speed inference.

Output defaults to:

build-cpu/

CUDA Build

Quick validation:

ED_BUILD_PROFILE=minimal bash scripts/build_cuda.sh

Performance validation:

CUDA_HOME=/path/to/cuda \
NCCL_ROOT=/path/to/nccl \
CUDNN_ROOT=/path/to/cudnn \
MPI_HOME=/path/to/mpi \
ED_BUILD_PROFILE=performance \
BUILD_DIR=build-cuda-performance \
bash scripts/build_cuda.sh

Output defaults to:

build-cuda/

Metal Build

Metal is macOS-only and experimental:

bash scripts/build_metal.sh

Output defaults to:

build-metal/

Vulkan Build

Vulkan is experimental. The helper script looks for glslc, Vulkan headers, and SPIRV-Headers:

bash scripts/build_vulkan.sh

Useful overrides:

VULKAN_SDK=/path/to/vulkan-sdk bash scripts/build_vulkan.sh
VK_EXTRA_HEADERS=/path/to/Vulkan-Headers/include \
SPIRV_HEADERS=/path/to/SPIRV-Headers/include \
bash scripts/build_vulkan.sh

Output defaults to:

build-vulkan/

Cleaning

Each build script accepts:

CLEAN=1 bash scripts/build_cpu.sh

CLEAN=1 only removes the selected build directory.

Common Build Errors

Missing ggml submodule

Symptom:

Missing ggml submodule at third_party/ggml

Fix:

bash scripts/bootstrap.sh

Missing NCCL in performance profile

Symptom:

ED_ENABLE_NCCL=ON requires NCCL headers and library

Fix:

NCCL_ROOT=/path/to/nccl ED_BUILD_PROFILE=performance bash scripts/build_cuda.sh

Missing MPI in performance profile

Install an MPI implementation or set:

MPI_HOME=/path/to/mpi

Missing cuDNN or cudnn-frontend

Install cuDNN and set:

CUDNN_ROOT=/path/to/cudnn

If the vendored cudnn-frontend source is absent, either initialize submodules or explicitly allow configure-time fetching:

ED_FETCH_CUDNN_FRONTEND=ON

Python package not found during tests

From the repository root:

PYTHONPATH=bindings/python/src python3 -m pytest bindings/python/tests

Related Documentation