An IREE compiler plugin for Coral NPU.
End-to-end flow for compiling and running JAX models (with plans to support other frontends) through IREE, targeting CoralNPU-style RISCV-32 execution. Currently, this uses host-side simulation instead of physical hardware.
To clone the project and its required submodules:
git clone sso://spacebeaker/coralnpu-compiler
cd coralnpu-compilerInitialize the top-level submodules and third_party/iree submodules in parallel (-j $(nproc)) with shallow depth (--depth=1), excluding the redundant duplicate checkout of llvm-project:
git submodule update --init --depth=1 -j $(nproc) -- .
git -C third_party/iree submodule update --init --depth=1 -j $(nproc) -- . ":(exclude)third_party/llvm-project"Note: We exclude third_party/llvm-project inside third_party/iree to avoid downloading a redundant ~2.4 GB duplicate copy of LLVM.
This is a temporary solution; we should have git forks of iree and llvm-project, that we can patch normally.
At the root of the project there are patches for third_party/iree and third_party/llvm-project.
Those can be applied using the script scripts/patch-third_party.sh:
./scripts/patch-third_party.sh --restore-first allSee the --help option for more details.
In any case, if you need to revert all the applied patches (uncommitted work will be lost):
git submodule foreach 'git clean -fd && git reset --hard HEAD'
git submodule update --init --depth=1 -j $(nproc) --force -- .
git -C third_party/iree submodule update --init --depth=1 -j $(nproc) --force -- . ":(exclude)third_party/llvm-project"We try to use bazel/cmake as much as possible to manage dependencies. The following are prerequisites that are not handled by bazel/cmake:
- git
- Bash >= 4.0
- Bazel 9.1.0
- clang 19
- lld 19
- cmake >= 3.26
- Python 3.13
- shfmt, for bash scripts formatting (https://github.com/mvdan/sh)
In a Debian based linux distro you can get all of the above like this:
sudo apt install git bash bazel-9.1.0 clang-19 lld-19 cmake python3.13 python3.13-venv shfmtTo install Bazel for other distributions, please refer to the official Bazel documentation.
In-tree dependencies are located in the third_party directory.
Iree requires a specific commit of llvm-project. We have it checked out in
third_party/llvm-project.
If you want to use a different revision of IREE, after checking it out in
third_party/iree, you can check the required llvm-project
commit hash by inspecting
third_party/iree/third_party/llvm-project
(e.g. git -C third_party/iree/third_party/llvm-project/ rev-parse HEAD), and
then checking it out in third_party/llvm-project.
Bazel's version matches third_party/coralnpu's.
bazel build --config=dev \
@iree_core//tools:iree-compile \
@iree_core//tools:iree-run-module \
@iree_core//compiler/bindings/python:compiler \
@iree_core//runtime/bindings/python:runtimeTo keep incremental builds as fast as possible, we use --fission=yes
(--config=dev does it, you don't need to do anything), which splits dwarf
information out of the .o files (see https://bazel.build/docs/user-manual). This
substantially reduces the input size to links and reduces link times
significantly.
Same as above, but use --config=release instead of --config=dev.
In general, you do not need to manually download or install Python packages (or Python), everything is managed through bazel. For some editors, you might need to recreate a similar environment as bazel's. You can do that like this:
python3.13 -m venv venv
. venv/bin/activate
pip install -r requirements_lock.txt
# if the above failes, try it with requirements.txtYou can also sandbox your Python 3.13 installation through conda/miniconda:
conda create --name venv python=3.13
conda activate venv
export PYTHONPATH=$CONDA_PREFIX
pip install -r requirements_lock.txtPut direct dependencies in requirements.txt, Run the following to update requirements_lock.txt
bazel run //:requirements.updateIf you use an LSP (e.g. clangd), you can run the following command to
generate/refresh compiler_commands.json:
bazel run --config=dev //:refresh_compile_commandsCreate a Python virtural environment with the required dependencies (See the Python section):
python3.13 -m venv venv
. venv/bin/activate
pip install -r requirements_lock.txt
# if the above failes, try it with requirements.txtSet BUILD_DIR to some directory where you want the build results to be.
For example BUILD_DIR=../coralnpu-compiler-build.
Run once:
cmake -G Ninja -B "${BUILD_DIR}" -S .Then, to build the compiler and runtime:
cmake --build "${BUILD_DIR}" --target iree-compile iree-run-module- Compiler is standalone: Building compiler targets (
iree-compile, IREE compiler plugins, LLVM/MLIR) via CMake is completely standalone and does not invoke or require Bazel. - Runtime Simulator Fallback: The simulator libraries are only needed by the runtime HAL driver. MPACT is always required; Spike and Verilator require
-DCORALNPU_ENABLE_SPIKE=ON/-DCORALNPU_ENABLE_VERILATOR=ON. Each is taken from-DCORALNPU_<NAME>_SIMULATOR_LIB=...if set, else frombazel-bin, else built from source with Bazel.
A normal compilation, without errors or warnings, does not print anything to stdout or stderr, unless a commandline option that specifically prints information is used.
# NB: anything before the -- will be interperted by bazel and not iree-compile
bazel run --config={dev|release} @iree_core//tools:iree-compile -- [iree-compile options]For example, to compile model.mlir:
# Compile for the host machine + CoralNPU (will run in simulation)
bazel run --config=dev @iree_core//tools:iree-compile -- \
--iree-hal-target-device=local \
--iree-hal-local-target-device-backends=llvm-cpu \
--iree-llvmcpu-target-cpu=host \
--iree-hal-target-device=coralnpu \
model.mlir \
-o model.vmfbSee the help message for the complete list of options:
bazel run --config=dev @iree_core//tools:iree-compile -- --helpCoralNPU compiler specific options are prefixed with --coralnpu.
BF16 accumulation: BF16 contractions, convolutions, and sum/product reductions accumulate in FP32 and round to BF16 once, on every device. This matches XLA but deviates from the StableHLO specification, which accumulates in the result type.
Affinity execution profile report:
--coralnpu-dump-affinity-profile-format={pretty|csv|json} dumps statistics about the compilation (such as the number of dispatches, estimated data size, and estimated work) grouped by the device affinity (e.g., host vs CoralNPU).
Register allocation report:
--coralnpu-dump-register-allocation-report-format={pretty|json}
--coralnpu-dump-register-allocation-report-dir=<directory>
--coralnpu-dump-register-allocation-report-filter=<pattern>
Dumps a report containing vector register utilization (unique registers used, vector configuration) and register allocator remarks (spills, reloads, copies) for each loop and at the function level. If the directory is -, the report is written to stdout; if empty or omitted, it is written to stderr. The filter option accepts a regex pattern to only report functions matching the name (default: .*dispatch.*|main).
Executable linking:
--coralnpu-link-executables={true|false} controls whether all executable dispatches are linked into a single library or emitted as individual self-contained executables (default: false). Emitting separate executables per dispatch avoids overflowing tightly constrained ITCM memory (e.g., 8 KB) on multi-dispatch models.
Affinity placement (roofline latency model):
--coralnpu-roofline-speedup-threshold=<ratio> (default: 1.0; 0 offloads all supported dispatches). A supported dispatch is placed on CoralNPU when T_cpu / T_npu >= threshold, otherwise on the host, where T_npu = T_launch + C_npu / B_copy + max(Ops / P_npu, Bytes / B_npu) and T_cpu = C_cpu / B_copy + max(Ops / P_cpu, Bytes / B_cpu). Ops and Bytes are taken over the whole dispatch and the narrowest input type of the root op picks the rate tier; C_npu / C_cpu are the input bytes produced on the other device, decided greedily in program order. Values not produced by a CoralNPU dispatch (looking through reshapes) count as host, except immutable weights and constants, which cost no copy. Without a host device, supported dispatches always go to CoralNPU.
Default constants (models a unified-memory SoC with an embedded Arm Cortex-A55 host, e.g. the Synaptics SL2619 Coralboard, where C_npu / C_cpu cross the m_axi link rather than the simulator's per-dispatch staging):
- CoralNPU:
P_npu= 128 GFLOPS (32-bit, Zvt 16x16 matrix unit), x2 at 16-bit, x4 at 8-bit;B_npu= 16 GB/s for the first 4 MiB (theEXTMEMwindow), 2 GB/s for bytes beyond it;T_launch= 800 ns;B_copy= 2 GB/s between host and CoralNPU (m_axi: 4-byte beats, 3.2 GB/s peak at 800 MHz). - Host CPU:
P_cpu= 8 GFLOPS (32-bit), x2 at 16-bit, x4 at 8-bit, but x1 for bf16 (no BF16 extension);B_cpu= 4 GB/s.
This repository supports building Python wheels, standalone binary distribution archives, and local installation trees via native Bazel targets.
To build and stage all release packages (dist tarball + Python wheels) into bazel-bin/output/:
bazel build --config=release //:outputThis generates:
bazel-bin/output/coralnpu-compiler-dist.tar.gzbazel-bin/output/coralnpu_compiler-0.0.1-py3-none-any.whlbazel-bin/output/coralnpu_runtime-0.0.1-py3-none-any.whl
Build release tarball containing bin/, lib/, crt/, and toolchain_rv32/ (saved under bazel-bin/build_tools/bazel/dist_tar.tar.gz):
bazel build --config=release //build_tools/bazel:dist_tarUnpack and install distribution archive directly to a specified directory:
bazel run --config=release //build_tools/bazel:install -- --prefix=/path/to/installTo verify that the installed compiler package and runtime binaries work end-to-end:
-
Save the following to
model.mlir:module { func.func @matmul(%arg0: tensor<32x64xf32>, %arg1: tensor<64x128xf32>) -> tensor<32x128xf32> { %cst = arith.constant 0.000000e+00 : f32 %0 = tensor.empty() : tensor<32x128xf32> %1 = linalg.fill ins(%cst : f32) outs(%0 : tensor<32x128xf32>) -> tensor<32x128xf32> %2 = linalg.matmul ins(%arg0, %arg1 : tensor<32x64xf32>, tensor<64x128xf32>) outs(%1 : tensor<32x128xf32>) -> tensor<32x128xf32> return %2 : tensor<32x128xf32> } } -
Compile an MLIR model targeting CoralNPU:
bazel run --config=dev @iree_core//tools:iree-compile -- \ --iree-hal-target-device=local \ --iree-hal-local-target-device-backends=llvm-cpu \ --iree-llvmcpu-target-cpu=host \ --iree-hal-target-device=coralnpu \ $(pwd)/model.mlir \ -o $(pwd)/model.vmfb -
Execute inference on the simulated CoralNPU device:
bazel run --config=dev @iree_core//tools:iree-run-module -- \ --device=coralnpu \ --module=$(pwd)/model.vmfb \ --function=matmul \ --input=32x64xf32=1.0 \ --input=64x128xf32=2.0
By default iree-run-module uses the MPACT functional simulator
(--simulator=mpact), which is always linked in. The Spike ISS, Verilator RTL
simulator, and FPGA hardware backend are optional and loaded from their shared
libraries on demand via LD_LIBRARY_PATH. Where --simulator cannot be
passed (e.g. the Python bindings), set CORALNPU_SIMULATOR=<name> instead; it
applies unless --simulator names another backend.
To run with Verilator (--simulator=verilator), build it once:
bazel build --config=dev @coralnpu_hw//hw_sim:libcoralnpu_simulator_vme.soand point LD_LIBRARY_PATH to its directory:
LD_LIBRARY_PATH=$(pwd)/bazel-bin/external/coralnpu_hw+/hw_sim \
bazel run --config=dev @iree_core//tools:iree-run-module -- \
--device=coralnpu \
--simulator=verilator \
--module=$(pwd)/model.vmfb \
...To run on FPGA hardware (--simulator=fpga), build the FPGA backend library:
bazel build --config=dev //runtime/sim/fpga:libcoralnpu_simulator_fpga.soand point LD_LIBRARY_PATH to its directory:
LD_LIBRARY_PATH=$(pwd)/bazel-bin/runtime/sim/fpga \
bazel run --config=dev @iree_core//tools:iree-run-module -- \
--device=coralnpu \
--simulator=fpga \
--module=$(pwd)/model.vmfb \
...Spike (--simulator=spike) works the same way with
//runtime/sim/spike:libcoralnpu_simulator_spike.so in bazel-bin/runtime/sim/spike.
For the CMake build, configure with -DCORALNPU_ENABLE_SPIKE=ON, -DCORALNPU_ENABLE_VERILATOR=ON, or -DCORALNPU_ENABLE_FPGA=ON instead; the libraries are copied to <build-dir>/runtime/sim, which still has to be on LD_LIBRARY_PATH.
To build Python wheels for the local host platform (saved under bazel-bin/build_tools/bazel/python_packages/...):
bazel build --config=release \
//build_tools/bazel/python_packages/coralnpu_compiler:wheel \
//build_tools/bazel/python_packages/coralnpu_runtime:wheelTo test the Python compiler (coralnpu_compiler) and runtime (coralnpu_runtime) wheel packages,
-
Create and activate a virtual environment:
python3.13 -m venv .venv source .venv/bin/activateThe Python version has to be 3.13, if this is not available see the Python section.
-
Build and install the Python wheels:
bazel build --config=release \ //build_tools/bazel/python_packages/coralnpu_compiler:wheel \ //build_tools/bazel/python_packages/coralnpu_runtime:wheelpip install \ bazel-bin/build_tools/bazel/python_packages/coralnpu_compiler/coralnpu_compiler-0.0.1-py3-none-any.whl \ bazel-bin/build_tools/bazel/python_packages/coralnpu_runtime/coralnpu_runtime-0.0.1-py3-none-any.whl -
Test if you can import and use the Python wheels:
python -c "import coralnpu.compiler as cnpuc; print(cnpuc)" <module 'coralnpu.compiler' from '/path/to/python3.13/site-packages/coralnpu/compiler/__init__.py'>
python -c "import coralnpu.runtime as cnpurt; print(cnpurt)" <module 'coralnpu.runtime' from '/path/to/python3.13/site-packages/coralnpu/runtime/__init__.py'>
-
Run end-to-end Python compilation and inference:
import numpy as np import coralnpu.compiler as coralnpu_compiler import coralnpu.runtime as coralnpu_runtime mlir_code = """ module { func.func @matmul(%arg0: tensor<32x64xf32>, %arg1: tensor<64x128xf32>) -> tensor<32x128xf32> { %cst = arith.constant 0.000000e+00 : f32 %0 = tensor.empty() : tensor<32x128xf32> %1 = linalg.fill ins(%cst : f32) outs(%0 : tensor<32x128xf32>) -> tensor<32x128xf32> %2 = linalg.matmul ins(%arg0, %arg1 : tensor<32x64xf32>, tensor<64x128xf32>) outs(%1 : tensor<32x128xf32>) -> tensor<32x128xf32> return %2 : tensor<32x128xf32> } } """ # Compile MLIR to VMFB bytes vmfb_bytes = coralnpu_compiler.compile_str( mlir_code, target_backends=["llvm-cpu", "coralnpu"], extra_args=[ "--iree-hal-target-device=local", "--iree-hal-local-target-device-backends=llvm-cpu", "--iree-llvmcpu-target-cpu=host", "--iree-hal-target-device=coralnpu", ], ) # Run inference on simulated CoralNPU config = coralnpu_runtime.Config("coralnpu") context = coralnpu_runtime.SystemContext(config=config) vm_module = coralnpu_runtime.VmModule.from_flatbuffer(context.instance, vmfb_bytes) context.add_vm_module(vm_module) arg0 = np.ones((32, 64), dtype=np.float32) arg1 = np.full((64, 128), 2.0, dtype=np.float32) result = context.modules.module.matmul(arg0, arg1) print("Output shape:", result.shape, "Output sample:", result[0, 0])
Build packages for all target platforms
bazel build --config=release //build_tools/bazel:all_platform_packagesTo cross-compile a single package for a specific target platform, pass --platforms=//build_tools/bazel/platforms:<platform>:
For example:
# Build distribution tarball for Linux AArch64 (ARM64)
bazel build --config=release --platforms=//build_tools/bazel/platforms:linux_aarch64 //build_tools/bazel:dist_tar//build_tools/bazel/platforms:linux_x86_64(Linux x86_64)//build_tools/bazel/platforms:linux_aarch64(Linux ARM64)//build_tools/bazel/platforms:macosx_x86_64(macOS Intel)//build_tools/bazel/platforms:macosx_arm64(macOS Apple Silicon)//build_tools/bazel/platforms:windows_x86_64(Windows x86_64)
Note
TODO (Cross-Compilation C++ Toolchains): The build system infrastructure (platform() targets, Starlark transitions, and wheel tagging) is in place for multi-platform builds. However, actually compiling C++ binaries for non-host platforms (e.g., linux_aarch64, macosx_arm64, windows_x86_64) requires registering corresponding C++ cross-compiler toolchains / sysroots (e.g. aarch64-linux-gnu, osxcross, mingw-w64) in MODULE.bazel. Currently, only the host C++ toolchain is registered.
Run all tests in the repository:
bazel test --config=dev //tests/...Run the CI test suite:
bazel test --config=dev //tests:ciWe have some StableHLO tests. To run just those:
bazel test --config=dev //tests/models/stablehlo/...We also have Linalg op tests. To run just those:
bazel test --config=dev //tests/models/linalg/...You can filter compiler tests by data-type tag (e.g., i8, i16, i32, f32):
# Run only f32 Linalg check tests
bazel test --config=dev //tests/models/linalg/... --test_tag_filters=f32
# Run only i8 StableHLO check tests
bazel test --config=dev //tests/models/stablehlo/... --test_tag_filters=i8Note
Running bazel test against //tests/models/linalg/... or //tests/models/stablehlo/... executes tests against static .mlir files checked into the repository under generated_<type>/ directories. Running tests will never trigger test regeneration.
The Linalg and StableHLO compiler check tests are generated from dynamic-shape templates using the check_gen tool and checked into the repository under generated_<type>/ directories (e.g., tests/models/linalg/generated_i8/).
When adding new test shapes or operations to op_tests_*.bzl, build and copy only the newly added tests:
bazel run --config=dev //tests:copy_missing
# or per-package:
# bazel run --config=dev //tests/models/linalg:copy_missing
# bazel run --config=dev //tests/models/stablehlo:copy_missingWhen removing or renaming test templates or instances:
bazel run --config=dev //tests:clean_stale
# or per-package:
# bazel run --config=dev //tests/models/linalg:clean_stale
# bazel run --config=dev //tests/models/stablehlo:clean_stalebazel run --config=dev //tests:sync_generatedVerify that all generated tests on disk match the declarations in op_tests_*.bzl:
bazel test --config=dev //tests:check_generatedForce reference evaluation and regenerate all tests across the repository. Delete all existing generated test files first to ensure no stale or obsolete tests remain:
rm -f tests/models/*/generated_*/*.mlir
bazel run --config=dev //tests:copy_generatedTests can be marked as manual (either on the op macro via tags = ["manual"] or on individual instance shapes via [("<shape>", ["manual"])]) when an operation or shape instance is work-in-progress or known to fail on CoralNPU:
- Always Generated: All declared tests in
op_tests_*.bzl(including manual tests) are generated intogenerated_<type>/bycopy_missing/copy_generated. - Excluded from CI: Tests tagged
manualare automatically skipped by//tests:ciand wildcard test runs (//tests/...). - Running a Manual Test: You can execute any individual manual test directly by its target label:
bazel test --config=dev //tests/models/linalg:<test_name>_<suffix>_check_test
To run the tests with CMake, you need to configure CMake with testing enabled (-DIREE_BUILD_TESTS=ON):
cmake -G Ninja -B "${BUILD_DIR}" -S . -DIREE_BUILD_TESTS=ONThen build the compiler, runtime, test runner, and test dependencies (including generated test bytecode modules):
# Build all test dependencies and test bytecode modules
cmake --build "${BUILD_DIR}" --target iree-test-deps -j $(nproc)
# (Optional) Or build only model test dependencies
cmake --build "${BUILD_DIR}" --target tests/models/all -j $(nproc)Finally, run the tests using ctest. Always pass -L "ci" to run the passing CI test suite and exclude manual/failing test instances:
# Run all CI tests
ctest --test-dir "${BUILD_DIR}" -L "ci" -j $(nproc)
# Run only StableHLO CI tests
ctest --test-dir "${BUILD_DIR}" -R "tests/models/stablehlo/.*" -L "ci" -j $(nproc)
# Run only Linalg CI tests
ctest --test-dir "${BUILD_DIR}" -R "tests/models/linalg/.*" -L "ci" -j $(nproc)You can also filter compiler check tests by data-type label (f32, i16, i32, i8):
# Run only f32 Linalg CI tests
ctest --test-dir "${BUILD_DIR}" -R "tests/models/linalg/.*" -L "f32" -LE "manual" -j $(nproc)
# Run only i8 StableHLO CI tests
ctest --test-dir "${BUILD_DIR}" -R "tests/models/stablehlo/.*" -L "i8" -LE "manual" -j $(nproc)Note
Like Bazel, running CMake check tests in tests/models/linalg/ or tests/models/stablehlo/ executes tests against static .mlir files checked into the repository under generated_<type>/ directories. Running tests with CMake will never trigger test regeneration. To regenerate check tests, use the Bazel copy_generated targets described above.
The compiler can be used to compile a JAX/StableHLO MLIR model to a vmfb binary that can be loaded by the IREE runtime Python bindings.
./examples/mobilenetv2-jax-aot/test_classify.shThe script first exports the model to mlir, using the StableHLO dialect. It then compiles the model to a vmfb, targeting the local host + CoralNPU. And finally runs an inference using the compiled model (the CoralNPU payload runs in simulation).
The toolchain supports compiling TFLite models via TOSA MLIR targeting CoralNPU:
./examples/mobilenetv2-tflite-aot/test_classify.shThe script:
- Downloads the official MobileNetV2 TFLite model.
- Legalizes the TFLite FlatBuffer to TOSA MLIR via
tosa-converter-for-tflite. - Compiles the TOSA model to a
.vmfbmodule targeting CoralNPU. - Executes image classification inference against the CoralNPU simulator.
The PJRT plugin invokes the IREE HAL device APIs and builds the dynamic library used by JAX for Just-in-Time (JIT) compilation and execution:
libiree_pjrt_coralnpu_dylib.so: Unified multi-device PJRT plugin supporting both CPU (device 0) and CoralNPU (device 1) targets.
./examples/gemma3-jax-pjrt/test_basic.shThis script:
- Builds
libIREECompiler.so,libiree_pjrt_coralnpu_dylib.so, andiree-compilevia Bazel. - Runs
examples/gemma3-jax-pjrt/basic.pyto test single-device CPU execution, single-device CoralNPU execution, and concurrent multi-device execution (CPU + CoralNPU) in JAX.
./examples/gemma3-jax-pjrt/test_chat.shThis script builds libIREECompiler.so, libiree_pjrt_coralnpu_dylib.so, iree-compile, and the coralnpu_tcm_highmem.ld linker script via Bazel, then runs end-to-end interactive multi-turn chat with Gemma3-270M across CPU + CoralNPU via JAX JIT compilation.
To run on CPU only, pass --cpu:
./examples/gemma3-jax-pjrt/test_chat.sh --cpuBy default only one token is generated per turn, since token generation is slow in the CoralNPU simulator. Pass --max_new_tokens to generate more (the CPU-only default is 128).
A tool to list all registered MLIR operations.
Using Bazel:
bazel run //tools/list-mlir-ops -- [dialect_namespace]Using CMake:
./$BUILD_DIR/tools/list-mlir-ops/list-mlir-ops [dialect_namespace]We use Google style, enforced by scripts/format-code.sh.
Before pushing anything, run the following command (NB: commit or stage your changes before, in case formatting does something horrible, and review the formatting changes).
scripts/format-code.shWe use clang 19, and lld (to build the compiler).
Places that need to be updated when changing version/toolchain:
Always use bash. Use this header:
#!/usr/bin/env bash
# Exit immediately on error (including in a pipeline), or when accessing an
# unset variable
set -euo pipefail