This notebook records my measurements of local LLM inference performance. I started with large mixture-of-experts models whose weights exceed GPU memory, placing the routed experts in system RAM while the dense computation runs on GPUs. By measuring where time goes in prompt processing and token generation, I can test changes to the configuration and to llama.cpp itself.
The most detailed investigation so far is GLM-5.2 on Galactus, an EPYC 7713 server with eight-channel DDR4 and four Radeon Pro V620 GPUs. Memory bandwidth explains much of its decode time. Prefill exposed a different problem: llama.cpp concentrated offloaded expert computation on one GPU, and a routing-index readback serialized the work. A scheduler patch raised prefill from 104.97 to 119.36 tokens/s (+13.7%) in the July comparison. Distributing the work alone produced no meaningful gain; removing the synchronization was also necessary.
The repository includes the measurements, unsuccessful experiments, patch instructions, and chronological notes behind these conclusions. The Galactus investigation now sits alongside the first oMLX measurements on Magneto, a 14-inch M2 Max MacBook Pro. Since the two engines were measured separately, each record gives its platform and conditions rather than treating the results as a common benchmark.
The latest common baseline in this notebook was measured on August 15–16, 2026, after Galactus's upgrade from 1 TB to 2 TB of RAM. All five models used llama.cpp build 3653e6d6d, the stock scheduler, CPU-resident routed experts, and -b 8192 -ub 8192 -p 8192 -n 128 -r 2. GLM used 32 threads; the other models used 64. pp8192 measures processing an 8,192-token prompt; tg128 measures generating 128 tokens. Rates are tokens per second.
| Model | pp8192 | tg128 | Speculative decode on the same build |
|---|---|---|---|
| MiniMax M2.7 | 418.83 ± 24.11 | 15.18 ± 0.18 | — |
| Qwen 3.5 397B | 249.61 ± 19.78 | 9.37 ± 0.16 | MTP failed to initialize on this export |
| DeepSeek-V4-Flash-0731 | 143.54 ± 1.64 | 10.34 ± 0.10 | 14.1 ± 0.7 (DSpark, n=3) |
| GLM-5.2 | 95.99 ± 3.36 | 5.30 ± 0.00 | 6.6 ± 0.3 (MTP, n=2) |
| Kimi K2.6 | 94.23 ± 4.45 | 5.79 ± 0.01 | — |
The speculative runs used a ZFS explanation prompt, greedy decoding, and llama-cli rather than llama-bench. Their repeated timings varied by about 9%; the reported spreads are not confidence intervals. MiniMax was retired after this measurement. Entry 12 records the exact model files, conditions, repetitions, and remaining questions.
The July patch result and April resident-expert results used different builds and configurations. They remain in the model notes, with their original conditions, and are not included in the stock baseline above.
Qwen3.8-27B on Magneto records an initial Lightning MTP baseline of 156.1 PP / 16.4 TG tokens/s at 4K, with tuned ANE + MTP at 170.1 / 13.9 and a later screenshot at 171.0 / 16.9. The baseline survives as a conversation transcription; the later results have a retained screenshot. Heavy thermal pressure was observed during the investigation. The oMLX leaderboard label omits chassis, so the higher external M2 Max result cannot establish an equivalent-machine comparison. A larger chassis is a hypothesis, not an identified cause. Entry 13 records the ANE tuner/benchmark distinction, DFlash regression, incomplete SpecPrefill experiment, and evidence limitations.
The findings collect what the experiments established and what remains uncertain. The methodology explains how I measured performance; the reproduction guide gives the commands and checks.
| Material | Where to find it |
|---|---|
| Model summaries and common baseline | Results |
| Scheduler investigation and exact edits | Prefill patch |
| Changes that did not help on Galactus | Negative results |
| MTP and DSpark measurements | Speculative decoding |
| Experiments in chronological order | Lab notebook |
| Early build record, dated model surveys, and deployment proposals | Source notes |
| Original captures and extracted measurements | Raw logs and CSV data |
| Machine specifications and memory measurements | Hardware |
| Diagnostic scripts and historical rerun protocol | Experiments |
The March build entry preserves the initial ROCm/container setup, DIMM replacement results, and expert-placement command, under conditions that differ from the July and August investigations. The dated source notes also retain a July model shortlist, September architecture comparisons, and an unmeasured two-node residency proposal. These records explain the build and candidate selection; they do not add a new common benchmark or establish the proposed routing roles.
As of the August 16 entry, the next comparison is the prefill patch against the stock build on the four retained models. An upstream patch submission, investigation of speculative-run timing variance, a choice between Kimi K2.6 and K2.7-Code, and a Kimi K3 evaluation remain open. The August speculative runs supersede the Session 10 rerun protocol.
Code in patches/ and experiments/ is covered by the MIT license. Text and data are covered by CC BY 4.0; LICENSING.md gives the scope and attribution details.
This project is not affiliated with AMD, llama.cpp, or any model vendor. Galactus and the other machine names are the names I use for the lab hardware.