A high-performance, lightweight, CPU-first LLM runtime foundation designed to become a drop-in alternative to llama.cpp.
Status: active inference-engine foundation. The project now includes llama.cpp-style CLI entry points, common llama.cpp flag parsing plus model-family-agnostic production-readiness validation including warning-clean MTP/speculative options with graceful fallback by default and strict mode, a minimal llama-compatible C ABI scaffold, streaming generation APIs, session/KV-cache management, sampler stack, real dense and MoE graph executors, RMSNorm/RoPE/attention/short-convolution/SwiGLU/logits operators, GGUF tokenizer loading with merge-aware BPE/trie fallback, real decode fail-fast boundary, internal q4_0 primitives, broad GGUF tensor type recognition, low-RAM mapped F16/Q4_0/Q4_1/Q8_0 matvec, hardened memory-mapped GGUF metadata/tensor-directory loading, tensor storage, mapped-weight benchmark scaffolding with finite deterministic inputs, warning-clean CPU feature detection, arena allocation, tokenizer scaffolding, model manifest loading, tests, and architecture notes.
- Drop-in llama.cpp ergonomics: build a
llama-clitarget and accept common flags like-m,-p,-n,--temp,-c,-t,--stream,--check,--dump-plan,--list-architectures,--accel,--spec-type mtp,--spec-type draft,--draft-model,--spec-draft-n-max, and--spec-strict. - Lightweight: small C++20 core, no required third-party dependencies.
- Fast by design: scratch arenas, zero-copy memory-mapped GGUF tensor access, real operator primitives, real MTP/speculative accept/reject core, q4_0 primitives, cache-aware kernels, runtime CPU feature dispatch, and a path toward fused graph execution.
- Novel architecture: clean separation between model format probing, tensor storage, session state, sampling, KV cache, and kernels.
- Production-oriented: testable layers, deterministic seeded sampling, streaming callback API, benchmark target, and compatibility wrappers at the edge.
- Apache-2.0: permissive licensing for research and production use.
apps/ cpullm-cli / llama-cli entry point
benchmarks/ microbenchmarks for kernel and mapped GGUF tensor work
include/cpullm/ public C++ API and llama-compatible C ABI scaffold
src/ runtime implementation
docs/ architecture, inference engine, compatibility plan, and roadmap notes
examples/ minimal embedding and toy manifest
tests/ dependency-free core tests
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j
ctest --test-dir build --output-on-failureEnable native CPU tuning for local benchmarking:
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release -DCPULLM_NATIVE=ON
cmake --build build -jNative name:
./build/cpullm-minimal
./build/cpullm-cli -m examples/toy-model.yml -p "Hello from cpullm" -n 32 --temp 0.8 --streamllama.cpp-style name:
./build/llama-cli -m examples/toy-model.yml -p "Hello from cpullm" -n 32 --temp 0.8 -c 2048 -t 0 --streamUniversal acceleration:
./build/cpullm-cli -m model.gguf --accel balanced -p "Hello"
./build/cpullm-cli -m model.gguf --accel turbo -p "Hello"Production readiness check:
./build/cpullm-cli -m model.gguf --check
./build/cpullm-cli -m model.gguf --dump-planBenchmark:
./build/cpullm-bench
./build/cpullm-gguf-q4-bench LFM2.5-230M-Q4_0.ggufcpullm::Arena— cache-line aligned scratch allocator for hot-path temporary memory.cpullm::detect_cpu_features()— portable runtime feature discovery for x86 and ARM targets.cpullm::matmul_f32()— blocked scalar baseline that future SIMD microkernels can replace through dispatch.cpullm::quantize_q4_0()/matvec_q4_0_f32()— compact internal quantized kernel primitives for lightweight inference.cpullm::matvec_gguf_any_f32()— low-RAM mapped GGUF matvec dispatch for F32, F16, Q4_0, Q4_1, and Q8_0 without full-weight dequantization.cpullm::dtype_name()/dtype_has_low_ram_matvec()— broad GGUF tensor type recognition for K-quants, IQ quants, TQ, BF16, and legacy quant formats.cpullm::Sampler— top-k, top-p, temperature, greedy mode, and deterministic seed support.cpullm::speculative_greedy_decode()— real callback-based draft/verify speculative accept/reject loop with stats; no mock token path.cpullm::mtp_greedy_decode()— MTP wrapper over the same production accept/reject core for built-in draft heads.cpullm::KVCache— explicit decoder cache capacity and memory accounting.cpullm::InferenceSession— reusable generation state with streaming token callbacks.cpullm::TensorStore— typed tensor registry for future memory-mapped model weights.cpullm::execute_dense_block()andexecute_moe_block()— real single-token dense/MoE graph executor foundations with top-k expert routing.cpullm::rms_norm,rope_inplace,causal_attention,short_convolution_1d,swiglu_mlp, andlogits_projection— real lightweight decoder operator primitives.cpullm::Tokenizer::from_gguf()— loads GGUF token arrays and merges, with cached token maps, merge BPE, trie longest-match fallback, and byte fallback.cpullm::greedy_decode_real()— real decode boundary that refuses unsupported GGUF architectures instead of mocking output.cpullm::architecture_profiles()— registry for common GGUF families including LLaMA, Qwen, Mistral/Mixtral, Gemma, Phi, DeepSeek, Granite, Falcon, MPT, StarCoder, LFM, and more.cpullm::build_execution_plan()— production-readiness planner for GGUF compatibility, kernel coverage, tensor requirements, and architecture blockers across model families.cpullm::apply_residency_policy()— universal GGUF tensor residency/prefetch acceleration that can use more RAM/startup compute to reduce cold page faults.cpullm::GgufFile— zero-copy memory-mapped GGUF metadata and tensor-directory loader with tensor byte views and robust numeric metadata extraction for LFM2/LFM2.5.cpullm::probe_gguf()— lightweight GGUF validation and metadata counts.cpullm::Tokenizer— minimal whitespace tokenizer placeholder for API development.cpullm::Model::load()— model loading front door for YAML manifests and GGUF probing.cpullm::Engine— generation facade ready for graph execution and real transformer layers.include/cpullm/llama_compat.h— minimal llama.cpp-style C ABI scaffold for migration experiments, routed through the sameModel::load()front door as the native CLI.
See docs/llama_compatibility.md, docs/mtp.md, docs/speculative_decoding.md, and docs/speculative_fallback.md.
Today, cpullm.cpp accepts common llama.cpp-style CLI invocations for YAML manifests, can memory-map GGUF metadata/tensor directories, and exposes MTP and classic draft-model speculative flags that gracefully fall back to normal decoding unless real verified speculative execution is possible; --spec-strict restores fail-fast behavior. Full LFM2.5 transformer operator execution is the next major milestone; the benchmark report now records llama.cpp as the real CPU baseline cpullm must beat.
After the real-operator/no-mock upgrade, llama.cpp b9821 runs LFM2.5-230M-Q4_0.gguf on the 2-vCPU AVX-512 sandbox at 519.19 ± 1.48 pp512 tokens/sec and 85.78 ± 0.88 tg128 tokens/sec. cpullm.cpp now loads/inspects the GGUF, loads tokenizer token arrays, runs real operator tests, and benchmarks a mapped GGUF Q4_0 tensor at 311.122 matvec/sec for blk.0.ffn_gate.weight, but it correctly refuses full LFM2.5 decode because the lfm2 graph is not wired yet. See docs/benchmarks/lfm25_230m_q4_cpu_after_real_ops.md.
The first real baseline report is available at docs/benchmarks/lfm25_230m_q4_cpu.md. On the 2-vCPU AVX-512 sandbox, llama.cpp b9821 runs LFM2.5-230M-Q4_0.gguf at 547.11 ± 2.01 pp512 tokens/sec and 87.54 ± 0.72 tg128 tokens/sec on CPU. cpullm.cpp is not fairly comparable yet because it can load mapped GGUF metadata/tensors but does not yet execute the full LFM2.5 transformer graph.
Reproduce with:
python3 scripts/benchmark_lfm25_230m_q4_cpu.py --threads 2See docs/tokenizer_parity.md, docs/inference_engine.md, docs/gguf_loader.md, docs/gguf_tensor_types.md, docs/real_executor.md, docs/production_readiness.md, docs/architecture_registry.md, and docs/universal_acceleration.md.
The engine is now structured around reusable sessions, explicit KV cache, quantized kernels, deterministic sampling, streaming callbacks, speculative fallback reporting, and zero-copy mapped GGUF tensor access. LFM2.5 metadata and Q4_0 tensors can be loaded and benchmarked directly from GGUF. GGUF generation now enters a real decode boundary and fails for unsupported architectures instead of producing synthetic text.
- Connect GGUF tensor mapping to the real dense/MoE executor foundations, then add family-specific LFM hybrid scheduling and quantized kernel dispatch, especially K-quant/IQ/TQ SIMD fused-dequant paths.
- Implement LFM2.5 RMSNorm, RoPE, attention, short-convolution/hybrid blocks, MLP, and logits projection operators.
- Add AVX2, AVX-512, and NEON q4_0/q8_0 microkernels behind runtime dispatch.
- Implement paged KV cache blocks and batch/decode scheduling.
- Wire MTP draft-head tensors and separate draft-model speculative decoding to the verifier/drafter executor for supported architectures.
- Broaden llama.cpp CLI and C API compatibility.
- Add reproducible benchmarks and profiling scripts against public model families.
Licensed under the Apache License, Version 2.0. See LICENSE.