cpullm.cpp now exposes classic draft-model speculative decoding as a real no-mock option.
./build/cpullm-cli \
-m target.gguf \
--draft-model draft.gguf \
--spec-type draft \
--spec-draft-n-max 4 \
-p "Hello"Compatibility aliases:
--spec-type speculative--spec-type draft-model--draft-model,--model-draft, or-md
Speculative decoding requires two real executors:
- a target/verifier model
- a smaller draft model
If the user requests speculative decoding before cpullm has a real verifier/draft executor wired for the loaded architecture, cpullm falls back to normal decoding by default and reports spec_active=false with a fallback reason. With --spec-strict, cpullm exits with an error. It never silently calls normal decoding speculative.
The repository includes the real greedy speculative accept/reject loop:
cpullm::speculative_greedy_decode(...)It is the same production algorithmic skeleton used by MTP:
- Generate a draft sequence from the draft model.
- Verify that sequence with the target model.
- Accept matching draft tokens.
- On first mismatch, emit the target token and continue from the corrected context.
- Track verifier steps, drafted tokens, accepted tokens, and rejected tokens.
The implementation is callback-based so it stays lightweight and can be wired to any future architecture executor without pulling in heavyweight dependencies.
- MTP uses draft heads inside the same GGUF/model.
- Classic speculative decoding uses a separate draft model.
Both share the same accept/reject core, but their draft callbacks come from different sources.