Patrick Buckley and Johannes Gäßler
db9d8aa428
ggml-cuda: native bf16 flash attention for vec kernel ( #20525 )
...
* ggml-cuda: native bf16 flash attention for vec and tile kernels
mma kernel still converts bf16 to fp16 before launch, native mma bf16 todo
* ggml-cuda: address code owner review feedback
reverted tile kernel changes to avoid larger refactor
* fix ci failures on turing and hip
* fix bf16 vec kernel compile on hip v_dot2 platforms
* add comments
---------
Co-authored-by: Johannes Gäßler <johannesg@5d6.de >
2026-03-22 11:05:51 +01:00
92f7da00b4
chore : correct typos [no ci] ( #20041 )
...
* fix(docs): correct typos found during code review
Non-functional changes only:
- Fixed minor spelling mistakes in comments
- Corrected typos in user-facing strings
- No variables, logic, or functional code was modified.
Signed-off-by: Marcel Petrick <mail@marcelpetrick.it >
* Update docs/backend/CANN.md
Co-authored-by: Aaron Teo <taronaeo@gmail.com >
* Revert "Auxiliary commit to revert individual files from 846d1c301281178efbc6ce6060ad34c1ebe45af8"
This reverts commit 02fcf0c7db661d5ff3eff96b2b2db9fdb7213256.
* Update tests/test-backend-ops.cpp
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
* Update tests/test-backend-ops.cpp
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
---------
Signed-off-by: Marcel Petrick <mail@marcelpetrick.it >
Co-authored-by: Aaron Teo <taronaeo@gmail.com >
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
2026-03-05 08:50:21 +01:00
Johannes Gäßler
5c662d21a3
CUDA: fix allignment on register spill for FA ( #18815 )
2026-01-15 15:14:50 +01:00
Daniel Bevenius
01cbdfd7eb
CUDA : fix typo in clang pragma comment [no ci] ( #18830 )
2026-01-14 10:31:49 +01:00
Johannes Gäßler
e95d0bc8fd
CUDA: fix FA VKQ accumulator overflow ( #17746 )
2025-12-05 09:18:10 +01:00
Johannes Gäßler and Aman Gupta <aman>
2e1c9cd814
CUDA: generalized (mma) FA, add Volta support ( #17505 )
...
* CUDA: generalized (mma) FA, add Volta support
* use struct for MMA FA kernel config
---------
Co-authored-by: Aman Gupta <aman>
2025-12-03 16:57:05 +01:00
R0CKSTAR and Johannes Gäßler
c6f7a423c8
[MUSA] enable fp16/fast_fp16/bf16_mma on PH1 ( #17551 )
...
* [MUSA] enable fp16/fast_fp16/bf16_mma on PH1
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com >
* Update ggml/src/ggml-cuda/fattn-vec.cuh
Co-authored-by: Johannes Gäßler <johannesg@5d6.de >
* Update ggml/src/ggml-cuda/fattn-vec.cuh
Co-authored-by: Johannes Gäßler <johannesg@5d6.de >
* Update ggml/src/ggml-cuda/fattn-tile.cuh
Co-authored-by: Johannes Gäßler <johannesg@5d6.de >
* Address review comments
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com >
---------
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com >
Co-authored-by: Johannes Gäßler <johannesg@5d6.de >
2025-11-28 14:08:29 +01:00
Johannes Gäßler
73955f7d2a
CUDA: no FP16 arithmetic for vector FA kernel ( #17558 )
2025-11-28 10:29:09 +01:00
Johannes Gäßler
9c7185dd28
CUDA: enable FA for FP32 KV cache ( #16546 )
2025-10-14 14:22:47 +02:00
R0CKSTAR
91a2a56556
musa: update compile flags ( #16265 )
...
Signed-off-by: Xiaodong Ye <yeahdongcn@gmail.com >
2025-10-02 16:29:56 +03:00
Johannes Gäßler
75a3a6c2cd
CUDA: refactor and deduplicate vector FA kernels ( #16208 )
...
* CUDA: refactor and deduplicate vector FA kernels
2025-09-27 18:45:07 +02:00