ggml: add ggml_backend_free_scratch to drop Vulkan compute prealloc
Add an optional backend interface method free_scratch (with a public ggml_backend_free_scratch wrapper) that frees transient/scratch device memory a backend holds outside of any allocated buffer, keeping the backend usable - the scratch is reallocated lazily on the next compute. Implement it for the Vulkan backend (ggml_backend_vk_free_scratch): free the prealloc_x/y/split_k/add_rms_partials and sync_staging device buffers and reset their sizes, so an idle/cold model does not hold the vision or matmul compute preallocations in VRAM. All other backends leave the hook null (no-op). Assisted-by: Claude
This commit is contained in:
@@ -431,6 +431,15 @@ void ggml_backend_synchronize(ggml_backend_t backend) {
|
||||
backend->iface.synchronize(backend);
|
||||
}
|
||||
|
||||
void ggml_backend_free_scratch(ggml_backend_t backend) {
|
||||
GGML_ASSERT(backend);
|
||||
if (backend->iface.free_scratch == NULL) {
|
||||
return;
|
||||
}
|
||||
|
||||
backend->iface.free_scratch(backend);
|
||||
}
|
||||
|
||||
ggml_backend_graph_plan_t ggml_backend_graph_plan_create(ggml_backend_t backend, struct ggml_cgraph * cgraph) {
|
||||
GGML_ASSERT(backend);
|
||||
GGML_ASSERT(backend->iface.graph_plan_create != NULL);
|
||||
|
||||
Reference in New Issue
Block a user