cpullm.cpp now has a production-oriented inference scaffold rather than a CLI-only placeholder.
- Model loading front door:
Model::load()dispatches between YAML manifests and GGUF probing. - GGUF probing: validates GGUF magic and reads version, tensor count, and metadata count without pulling the whole file into memory.
- Tensor store: central typed tensor registry for future memory-mapped model weights.
- Quantized kernels: q4_0 packing plus q4_0 x f32 dot/matvec primitives.
- Sampler: temperature, top-k, top-p, deterministic seedable PRNG, and greedy mode.
- KV cache: explicit capacity and allocation accounting for decoder sessions.
- Session API: reusable
InferenceSessionwith string generation and streaming token callbacks. - CLI streaming:
--streamemits token events immediately. - Benchmark harness:
cpullm-benchgives a simple q4_0 matvec timing target.
The engine is intentionally split into small independently testable pieces. Real model execution will connect GGUF tensor loading to transformer graph execution through this path:
- Memory-map GGUF tensor data.
- Register tensors in
TensorStore. - Select quantized kernels through CPU feature dispatch.
- Execute decoder layers with paged KV-cache slots.
- Stream sampler output through
InferenceSession.
This is a clean-room path toward llama.cpp-style ergonomics with a smaller core and sharper extension points.