Kernel is a port of `ggml-cuda/fwht.cu` (us/run, median): ``` m x n x k GEMM FWHT speedup 64 x 1 x 64 10.20 2.93 3.48x 64 x 2048 x 64 10.75 2.71 3.97x 128 x 1 x 128 10.33 2.88 3.59x 128 x 32 x 128 9.20 2.77 3.33x 128 x 2048 x 128 16.46 2.76 5.95x 256 x 1 x 256 10.19 2.77 3.68x 256 x 2048 x 256 16.69 3.41 4.89x 512 x 2048 x 512 54.16 12.89 4.20x ```
13 lines
492 B
C++
13 lines
492 B
C++
#ifndef GGML_SYCL_FWHT_HPP
|
|
#define GGML_SYCL_FWHT_HPP
|
|
|
|
#include "common.hpp"
|
|
|
|
// Fast Walsh-Hadamard transform, the fast path for a MUL_MAT whose src0 ggml has
|
|
// tagged GGML_HINT_SRC0_IS_HADAMARD. src0 is not read at all. Returns false if the
|
|
// shape is not one this can serve, in which case the caller must fall through to the
|
|
// ordinary mat-mul dispatch.
|
|
bool ggml_sycl_op_fwht(ggml_backend_sycl_context & ctx, const ggml_tensor * src, ggml_tensor * dst);
|
|
|
|
#endif // GGML_SYCL_FWHT_HPP
|