server: optionally evict the KV cache too (Phase 2 of VRAM sharing)
Extend on-demand device residency to the KV cache so that when a model's KV plus another model would not fit in VRAM, the KV can also be evicted to a host shadow (D2H on release, H2D on restore) instead of only the weights. - llama_memory_i: add release_device_buffers()/restore_device_buffers() (default no-op). Implemented in llama_kv_cache (D2H shadow of the live ctxs_bufs, freed and reallocated like the weights); llama_memory_hybrid and llama_kv_cache_iswa delegate to their child caches. - llama_context::release_device(evict_kv): also evict the memory's device buffers when requested; restore_device() rebuilds them. Public API llama_context_release_device gains an evict_kv flag. - server: LLAMA_SLEEP_EVICT_KV=1 enables it. Off by default (weights-only), since the KV shadow adds a D2H/H2D copy of the live cache each cycle. Validated on RX 580 (Vulkan), 4B @ 32k ctx: weights-only cold VRAM 1750 MB (KV stays); weights+KV cold VRAM 726 MB (KV freed, ~1 GB reclaimed). KV survives the round-trip: prompt cache reused after the cycle (prompt_n 4 vs 42), correct output. Assisted-by: Claude
This commit is contained in:
+7
-2
@@ -59,8 +59,10 @@ struct llama_context {
|
||||
|
||||
// on-demand device (VRAM) residency: free / rebuild the model's GPU weight buffers so that
|
||||
// VRAM can be time-shared with another model. release_device() synchronizes first; decode()
|
||||
// auto-restores when it finds the weights have been released.
|
||||
void release_device();
|
||||
// auto-restores when it finds the weights have been released. If evict_kv is set, the KV cache
|
||||
// device buffers are also evicted to a host shadow (for when the KV would not fit alongside the
|
||||
// other model) - this adds a D2H/H2D copy of the live KV on each cycle.
|
||||
void release_device(bool evict_kv = false);
|
||||
void restore_device();
|
||||
|
||||
const llama_model & get_model() const;
|
||||
@@ -351,6 +353,9 @@ private:
|
||||
|
||||
bool sched_need_reserve = true;
|
||||
|
||||
// on-demand device residency: true while the KV cache device buffers have been evicted to host
|
||||
bool kv_device_evicted = false;
|
||||
|
||||
ggml_backend_t backend_cpu = nullptr;
|
||||
std::vector<ggml_backend_ptr> backends;
|
||||
|
||||
|
||||
Reference in New Issue
Block a user