server: optionally evict the KV cache too (Phase 2 of VRAM sharing)

Extend on-demand device residency to the KV cache so that when a model's KV
plus another model would not fit in VRAM, the KV can also be evicted to a host
shadow (D2H on release, H2D on restore) instead of only the weights.

- llama_memory_i: add release_device_buffers()/restore_device_buffers()
  (default no-op). Implemented in llama_kv_cache (D2H shadow of the live
  ctxs_bufs, freed and reallocated like the weights); llama_memory_hybrid and
  llama_kv_cache_iswa delegate to their child caches.
- llama_context::release_device(evict_kv): also evict the memory's device
  buffers when requested; restore_device() rebuilds them. Public API
  llama_context_release_device gains an evict_kv flag.
- server: LLAMA_SLEEP_EVICT_KV=1 enables it. Off by default (weights-only),
  since the KV shadow adds a D2H/H2D copy of the live cache each cycle.

Validated on RX 580 (Vulkan), 4B @ 32k ctx: weights-only cold VRAM 1750 MB
(KV stays); weights+KV cold VRAM 726 MB (KV freed, ~1 GB reclaimed). KV
survives the round-trip: prompt cache reused after the cycle (prompt_n 4 vs
42), correct output.

Assisted-by: Claude
This commit is contained in:
2026-07-26 21:58:43 +02:00
parent 0b0cc69635
commit 62873a34bb
11 changed files with 156 additions and 12 deletions
+4 -2
View File
@@ -568,9 +568,11 @@ extern "C" {
// On-demand device (VRAM) residency: free the model's GPU weight buffers (keeping a host
// shadow so no reload is needed) to hand VRAM to another model, and rebuild them on demand.
// The KV cache is left resident. Any subsequent llama_decode auto-restores the weights.
// Any subsequent llama_decode auto-restores. If evict_kv is true, the KV cache device buffers
// are also evicted to a host shadow (for when the KV would not fit alongside the other model),
// at the cost of a D2H/H2D copy of the live KV each cycle; otherwise the KV stays resident.
// Intended for time-sharing a single GPU between several always-loaded models.
LLAMA_API void llama_context_release_device(struct llama_context * ctx);
LLAMA_API void llama_context_release_device(struct llama_context * ctx, bool evict_kv);
LLAMA_API void llama_context_restore_device(struct llama_context * ctx);
LLAMA_API bool llama_context_weights_resident(const struct llama_context * ctx);
LLAMA_API enum llama_pooling_type llama_pooling_type(const struct llama_context * ctx); // TODO: rename to llama_get_pooling_type