Skip to content

Latest commit

 

History

History
54 lines (36 loc) · 1.71 KB

File metadata and controls

54 lines (36 loc) · 1.71 KB

Speculative Decoding

cpullm.cpp now exposes classic draft-model speculative decoding as a real no-mock option.

CLI

./build/cpullm-cli \
  -m target.gguf \
  --draft-model draft.gguf \
  --spec-type draft \
  --spec-draft-n-max 4 \
  -p "Hello"

Compatibility aliases:

  • --spec-type speculative
  • --spec-type draft-model
  • --draft-model, --model-draft, or -md

No mock policy

Speculative decoding requires two real executors:

  1. a target/verifier model
  2. a smaller draft model

If the user requests speculative decoding before cpullm has a real verifier/draft executor wired for the loaded architecture, cpullm falls back to normal decoding by default and reports spec_active=false with a fallback reason. With --spec-strict, cpullm exits with an error. It never silently calls normal decoding speculative.

Implemented core

The repository includes the real greedy speculative accept/reject loop:

cpullm::speculative_greedy_decode(...)

It is the same production algorithmic skeleton used by MTP:

  1. Generate a draft sequence from the draft model.
  2. Verify that sequence with the target model.
  3. Accept matching draft tokens.
  4. On first mismatch, emit the target token and continue from the corrected context.
  5. Track verifier steps, drafted tokens, accepted tokens, and rejected tokens.

The implementation is callback-based so it stays lightweight and can be wired to any future architecture executor without pulling in heavyweight dependencies.

Difference from MTP

  • MTP uses draft heads inside the same GGUF/model.
  • Classic speculative decoding uses a separate draft model.

Both share the same accept/reject core, but their draft callbacks come from different sources.