You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit d37d555
Browse filesBrowse the repository at this point in the historyBrowse files
print("[SwiftLM] ⚠️ --turbo-kv has no effect with --ctx-size \(ctx): a bounded context gives the attention layers a RotatingKVCache, and TurboKV only compresses KVCacheSimple. Drop --ctx-size to use --turbo-kv.")
// This compresses cache history older than 8192 tokens into 3.5-bit Polar+QJL
2039
2047
// form, halving KV RAM for long-context (100k+) requests.
2040
2048
if config.turboKV {
2049
+
varenabledLayers=0
2041
2050
forlayerin cache {
2042
2051
iflet simple = layer as?KVCacheSimple{
2043
2052
simple.turboQuantEnabled =true
2053
+
enabledLayers +=1
2044
2054
}
2045
2055
}
2056
+
if enabledLayers ==0 && !TurboKVNotice.warnedNoLayers {
2057
+
// Warn once. A benign race here only means a duplicate log line.
2058
+
TurboKVNotice.warnedNoLayers =true
2059
+
print("[SwiftLM] ⚠️ --turbo-kv is enabled but this model's cache has no KVCacheSimple layers (\(cache.count) layers, e.g. RotatingKVCache from --ctx-size), so no KV compression is applied.")
2060
+
}
2046
2061
}
2047
2062
2048
2063
// ── Prompt cache: bypass for multimodal inputs ──
0 commit comments