This website requires JavaScript.
e107984bcf
ops: add Hexagon to ops.md and update main README.md (#28263 )
Todor Boinovski
2026-09-03 07:14:36 -07:00
42f0225fea
server : use pytest-xdist for server tests (#28298 )
Daniel Bevenius
2026-09-03 15:04:30 +02:00
de8656bd94
mtmd: propagate const to preproc class (#28310 )
Xuan-Son Nguyen
2026-09-03 12:57:10 +02:00
7bb0fc18f6
metal : add sparse FA (#28098 )
Georgi Gerganov
2026-09-03 13:51:13 +03:00
0df017d6dd
metal : fix glu dispatch with ne00 = 1 (#28306 )
Georgi Gerganov
2026-09-03 13:25:41 +03:00
f45576aa86
mtmd : add const in various places (#28307 )
Mads Marquart
2026-09-03 12:12:49 +02:00
0ba6499c3b
CUDA: Allow concurrent streams per split for multi-GPU (#28198 )
2026-09-03 12:03:03 +02:00
c7bda030e7
vulkan: fix FA dequant path engagement (#28190 )
Nathan Wilson
2026-09-03 15:40:34 +07:00
0df974d777
sycl : enhance the api to support peer-to-peer copy (#27550 )
Neo Zhang
2026-09-03 15:41:07 +08:00
d646c9d155
convert : skip bias_vl tensor in DeepSeek-V4 DSpark conversion (#28294 )
Georgi Gerganov and Sigbjørn Skjæret
2026-09-03 10:37:23 +03:00
5ec4eab69e
misc : prevent RAM peaking at model loading stage (#27483 )
Tarek Dakhran
2026-09-03 09:32:24 +02:00
4aa6ffba25
sycl: reduce redundant work in Q4_K multi-column MMVQ (#27062 )
Eurekatic and RaulAbejonDelgado
2026-09-03 08:59:06 +02:00
c61b98b875
model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444 )
Yaniss Amazouz
2026-09-03 09:53:08 +03:00
67a17c17ca
mtmd: fix idefics3 preproc (#28273 )
Xuan-Son Nguyen
2026-09-03 01:00:57 +02:00
159b741427
finetune: fix no KV cache (#27199 )
Xuan-Son Nguyen
2026-09-02 23:53:32 +02:00
9cffdcc801
server : accept data: URLs for input_video and input_audio (#27735 )
Abhiram
2026-09-03 01:54:31 +05:30
f027c4f1b0
ggml-hexagon: add F16 support for unary ops (#28228 )
cqderek
2026-09-03 03:59:36 +08:00
7339054744
mtmd: add mtmd_tokenize_from_parts() (#28250 )
Xuan-Son Nguyen
2026-09-02 21:20:10 +02:00
9cc33944f9
metal : add fa-vec tunings for M3 (#28236 )
Isaac
2026-09-02 23:43:12 +05:30
8c0b9cd04a
metal : fix memory query under low-memory conditions (#27701 )
2026-09-02 20:09:46 +02:00
03dbcc53e1
ci : check for missing autoreleasepools (#27884 )
Niklas Wenzel
2026-09-02 19:54:48 +02:00
cff184438e
Update ROCm to 10.0.0 release (#27803 )
Mario Limonciello
2026-09-02 12:49:11 -05:00
9400c8946e
model: correctly support input vision for deepseek4 (#28154 )
Xuan-Son Nguyen
2026-09-02 19:14:46 +02:00
d5fec32a87
ci : enable hf-jobs on server-cuda (#28258 )
Sigbjørn Skjæret
2026-09-02 19:13:20 +02:00
3d3d7c8181
ggml-cuda : remove unused vars (#28235 )
Adrien Gallouët
2026-09-02 18:54:11 +02:00
e750b887a8
common, server : enable preserve_reasoning kwarg by default, log its effective state (#28174 )
Georgi Gerganov and Xuan-Son Nguyen
2026-09-02 19:19:54 +03:00
7798007a29
mtmd: support DeepSeek-V4-Flash-Vision-Exp (#28133 )
Xuan-Son Nguyen
2026-09-02 16:43:43 +02:00
8e93a9773b
CUDA + ggml: add sparse-fa for DSV4/GLM (#27970 )
Aman Gupta
2026-09-02 19:57:37 +05:30
0f3a71be15
mtmd: Fix Qwen3-tts-0.6b (#28231 )
Pascal
2026-09-02 12:46:16 +02:00
b81c99b479
ggml: avoid KleidiAI buffer type init on dispatch (#27891 )
Aman Chadha(IVIXMMI) and Acmmi
2026-09-02 11:46:15 +05:30
960dffab05
hexagon: MUL_MAT and MUL_MAT_ID fusion and fixes (#28202 )
Max Krasnyansky
2026-09-01 23:15:21 -07:00
ba8818cbf3
vulkan: handle larger batch sizes (>4) efficiently for IQ3_S mat-vec (#27449 )
Laurent Zuijdwijk and Marshall
2026-09-02 07:14:52 +01:00
56dd8150cc
vulkan : only request VK_KHR_shader_bfloat16 extension if supported (#28155 )
Mads Marquart
2026-09-02 08:13:25 +02:00
2637dfe373
ggml-cpu : conditionally add SpacemiT IME kernel sources (#27961 )
Alan Tseng
2026-09-02 14:12:28 +08:00
43d87ff2dd
opencl: fix out‐of‐bound reads in the Adreno image kernels (#27632 )
Hongqiang Wang
2026-09-01 22:28:45 -07:00
69320fef12
hexagon: add missing FARF logs for cpy/get_rows/set_rows/gdn ops (#28217 )
Trivikram Reddy
2026-09-02 00:20:29 -05:00
b96806d960
metal : add metallib build support for xcframework (#28163 )
Jhen-Jie Hong
2026-09-02 07:45:56 +08:00
3466812d1f
cuda: fuse MoE weighted expert reduction (#25952 )
anujj
2026-09-02 01:18:47 +05:30
b356fa2624
kv-cells: look up the n-gram history in the sequence position index (#28040 )
Pascal
2026-09-01 20:16:07 +02:00
dfc29b64eb
context : autoscale n_ctx_train when yarn scaling specified (#28030 )
Sigbjørn Skjæret
2026-09-01 18:59:54 +02:00
f28493c783
models : appropriately flag noscan ssm_a tensors (#28121 )
Sigbjørn Skjæret
2026-09-01 18:59:15 +02:00
73159c3039
model : fix gemma4-assistant (#28183 )
Sigbjørn Skjæret
2026-09-01 18:58:44 +02:00
d11b3cc7ed
model : load relevant arrays with n_layer_all (#28173 )
Sigbjørn Skjæret
2026-09-01 18:58:29 +02:00
c845263f8b
Revert "sycl : add Kronecker product FWHT support for sizes 384, 640, 768, 12…" (#28184 )
Titaniumtown
2026-09-01 09:04:31 -07:00
1f3d318734
sycl : add Kronecker product FWHT support for sizes 384, 640, 768, 1280 (#28016 )
Jingxin (Philip) Li
2026-09-01 23:47:08 +08:00
8887a48f05
metal : add fa-vec tuning for M2 Pro (#28122 )
Lukasz Stolcman
2026-09-01 15:24:44 +02:00
be789c3448
metal : add fa-vec tunings for A18 Pro (MacBook Neo) (#28152 )
Jhen-Jie Hong
2026-09-01 21:15:59 +08:00
9d817213a0
model : load hparams.n_layer_nextn before n_layer() calls (#28159 )
Sigbjørn Skjæret
2026-09-01 13:55:45 +02:00
fe2120bc9d
metal : fix more leaks due to missing autoreleasepools (#27883 )
Niklas Wenzel and YiChen Lv
2026-09-01 13:50:47 +02:00
d08c7872d6
metal : add fa-vec tuning for M2 Max (#28015 )
Georgi Gerganov
2026-09-01 13:37:40 +03:00
5eec3ad017
sycl : support limit max alloc memory within 2GB for host-pinned memory (#27559 )
Neo Zhang
2026-09-01 18:35:47 +08:00
36b1015438
qwen4exp: fix seq_cp, block position keying, mtmd input, cuda abort, add tests (#27941 )
Daniel Han
2026-09-01 03:22:04 -07:00
d086dbb348
tests : fix log verbosity for test-llama-archs (#28147 )
Georgi Gerganov
2026-09-01 13:07:12 +03:00
1b89a43e38
quantize: row-slab stream to avoid thread starvation (#27830 )
Xuan-Son Nguyen
2026-09-01 11:18:54 +02:00
d5d993a093
metal: enable Metal 4.0 tensor API on M5+/A19+ (#27461 )
James Francis
2026-09-01 03:02:42 -06:00
234a6ebaa0
ci: Bump ggml-org/ccache-action to v1.2.24 (#28083 )
Ludovic Henry
2026-09-01 11:00:17 +02:00
518b76236b
kleidiai : Update KleidiAI Documentation (#26078 )
Jonathan Clohessy
2026-09-01 09:45:13 +01:00
0eadefebd3
qwen4exp: support recurrent state rollback (#28123 )
Pascal
2026-09-01 06:24:49 +02:00
09412af38a
qwen4exp: sum the indexer heads by slices (#28023 )
Pascal
2026-09-01 06:23:59 +02:00
458681e1d5
metal : add fa-vec tunings for M1 Ultra (#28088 )
Buğra Özgürsoy
2026-09-01 00:47:27 +03:00
e4b9af007b
CUDA: XOR swizzle flash attn K,V smem fp16 tiles (#25635 )
ynankani
2026-08-31 20:18:01 +00:00
ab0b3bd3c8
metal : add concat support for quantized types (#28116 )
Georgi Gerganov
2026-08-31 23:16:04 +03:00
85c55223ca
AVX2: Speed up large batch size prompt processing of IQ models (#27402 )
Bartowski and Georgi Gerganov
2026-08-31 14:33:50 -04:00
2a74817f93
metal : add top-k radix implementation (#28073 )
Georgi Gerganov
2026-08-31 21:31:53 +03:00
2d8d612e4c
kv-cache : optimize restoring non-contiguous cells (#27991 )
itsnotoger
2026-08-31 18:49:58 +02:00
010be9683a
opencl: tune the quant paths for Intel Xe-LP GPUs to improve its TG and PP performance (#26438 )
Hongqiang Wang
2026-08-31 08:56:22 -07:00
774ee0e200
ui: copy the displayed text of grouped agentic responses (#27832 )
Pascal
2026-08-31 17:48:43 +02:00
8e53fcefd2
webgpu : avoid crash when offset is not multiple of 4 in WebGPU ggml_backend_tensor_get() implementation (#28045 )
2026-08-31 16:04:38 +02:00
f8dbcd6189
ROCm: add radix TOP_K for long rows (#27466 )
Jaden_Mach
2026-08-31 09:00:04 -04:00
5d4a3be26d
metal : add fa-vec tunings for M1 (#28078 )
Niklas Wenzel
2026-08-31 13:58:55 +02:00
41ef91f7c8
CUDA: extend MOE fusion to specdec, earlier MOE glu fusion and topk-router fusion were restricted to 1 token (#27621 )
ynankani
2026-08-31 11:22:28 +00:00
a32af33de2
sycl : Enhance to get the free memory of Intel GPU (#27968 )
Neo Zhang
2026-08-31 18:33:02 +08:00
580e88d8b7
ci : add check for unzip (#28082 )
Sigbjørn Skjæret
2026-08-31 12:17:51 +02:00
662a0b0121
spec : fuse the DFlash encoder into the KV cache injection (#27310 )
2026-08-31 17:19:20 +08:00
2cdae802e4
vulkan: tune mat-vec rows for batched inference on Strix Halo (#27909 )
Simon Teixidor
2026-08-31 11:07:53 +02:00
557614e029
ggml : add MUL_MAT to the list of ops that may need additional memory (for WebGPU) (#28071 )
fairydreaming and Stanisław Szymczyk
2026-08-31 10:17:23 +02:00
daef7b6874
vulkan: top_k radix select for k >= 1024 for Qwen 3.8 Flash Next (#28032 )
Ruben Ortlam
2026-08-31 07:04:34 +02:00
9723942adc
hexagon: fix CPY fence bug (#28033 )
Shenghan Yang
2026-08-31 02:18:24 +08:00
bd55e6aae8
metal : add remaining Q4_1/Q5_0/Q5_1 fa-vec tunings for M2 (#28017 )
codemonkey
2026-08-31 02:00:10 +08:00
a7cc83bbae
rpc: avoid serializing buffers from other servers (#26500 )
hmirin and Georgi Gerganov
2026-08-31 02:26:16 +09:00
6d1479c148
ggml : fix ggml_backend_buft_get_alloc_size() guard (#28038 )
Georgi Gerganov
2026-08-30 20:25:15 +03:00
62acc89c26
kv-cells: stop the sequence scan once all sequences are seen (#28011 )
Pascal
2026-08-30 17:27:34 +02:00
0190529ec4
ggml: add SWIGLU_CLAMP (#27930 )
Aman Gupta
2026-08-30 20:30:02 +05:30
2578138397
llama: improve TENSOR_READ_LAZY handling (#27837 )
Xuan-Son Nguyen
2026-08-30 16:59:48 +02:00
f1793c1c4e
CUDA: use the fast mm_ids_helper path for any n_expert_used (#27978 )
Pascal
2026-08-30 16:06:32 +02:00
0b5be7e4a2
hip: tune rdna 3 mmq config (#26284 )
itterative
2026-08-30 13:47:21 +03:00
e422148047
hip : optimize Q2_0 dot-product path for gfx1201 (#26753 )
LunalFresh
2026-08-30 05:18:36 -05:00
cc231cb0da
dflash: pass missing NVFP4 scales to attention operations (#28000 )
JamePeng
2026-08-30 16:34:39 +08:00
bebc9350ec
common: rename --tensor-read-lazy to --lazy-mode, add -lzm shorthand (#27969 )
Georgi Gerganov
2026-08-30 09:18:10 +03:00
73f56d105b
ggml : add ggml_backend_op_alloc_size_may_expand, use it in RPC (#27960 )
Georgi Gerganov
2026-08-30 09:17:47 +03:00
742347b2e7
rpc: fix apple rdma error spew on teardown (#27908 )
Ryan C
2026-08-30 06:16:26 +00:00
093adb242e
metal: add fa-vec tunings for M3 Ultra (#27999 )
Nils Gladitz
2026-08-30 08:06:29 +02:00
b8b743c3c1
metal : Add fa-vec tuning for M3 Pro (#27963 )
Daya Adianto
2026-08-30 06:02:22 +00:00
dc7aecf70d
vendor : update cpp-httplib to 0.54.0 (#27919 )
Alessandro de Oliveira Faria (A.K.A.CABELO)
2026-08-30 03:01:51 -03:00
2bf0415152
rpc : fix pre-rdma macOS versions (#27815 )
Ryan C
2026-08-30 05:59:25 +00:00
9e54e687cb
hexagon: support for device discovery and create sessions on demand (#27785 )
2026-08-29 22:57:55 -07:00
370cb12e8b
sycl: split long rows in TOP_K instead of one work-group per row (#27847 )
Titaniumtown
2026-08-29 22:57:08 -07:00
d882575cc8
metal : fix null-pipeline crash for F16 src1 mul_mat/mul_mat_id (#25648 )
QuintinShaw
2026-08-30 13:56:35 +08:00
bdf3955159
memory : copy Hadamard matrix to k_rot tensor only if it has buffer assigned to prevent crashes during context shift of unquantized K cache (#27967 )
2026-08-30 07:47:15 +02:00
57291f2644
ggml: allow passing alloc dependencies in graph_optimize (#27301 )
Aman Gupta
2026-08-30 09:04:20 +05:30