CUDA: use mmvq for mul-mat-id for small batch sizes (#18958)

* CUDA: use mmvq for mul-mat-id for small batch sizes

* add mmvq too

* Fix perf issue on ampere. Use mmvf mm-id only for non-nvidia GPUs

* templatize multi_token_path
This commit is contained in:
Aman Gupta
2026-02-03 23:31:23 +08:00
committed by GitHub
parent a6fd8ca1fe
commit 8bece2eb20
4 changed files with 224 additions and 121 deletions
+2
View File
@@ -1,5 +1,7 @@
#include "common.cuh"
#define MMVF_MAX_BATCH_SIZE 8 // Max. batch size for which to use MMVF kernels.
void ggml_cuda_mul_mat_vec_f(ggml_backend_cuda_context & ctx, const ggml_tensor * src0, const ggml_tensor * src1, const ggml_tensor * ids, ggml_tensor * dst,
const ggml_cuda_mm_fusion_args_host * fusion = nullptr);