This website requires JavaScript.
b4aa7dd477
mtmd : use align_corners for qwen3vl vision position embedding interpolation (#25781 )
Gerben van V and GitHub
2026-07-21 23:58:34 +02:00
71102a73f2
hexagon: check tensor type when reusing descriptors (#25968 )
Wei Wang and GitHub
2026-07-22 05:44:22 +08:00
846e991ec3
cuda: add sqrt_softplus in topk-moe for dsv4 (#25896 )
Aman Gupta and GitHub
2026-07-22 00:30:01 +08:00
fb0e6b6219
kleidiai : warn once when a weight type has no KleidiAI kernel (#25701 )
Kamalesh VS and GitHub
2026-07-21 21:40:29 +05:30
60f6a17704
common: resolve draft repo to its requested sidecar (#25955 )
Pascal and GitHub
2026-07-21 18:03:43 +02:00
fd41bf65a2
server: return 400 instead of 500 on validation error with X-Conversation-Id (#25760 )
Pascal and GitHub
2026-07-21 17:47:54 +02:00
40b740ad05
server : properly handle null llama_context (#25868 )
2026-07-21 17:47:17 +02:00
f048010180
vulkan: Refactor vk_queue to use per-instance mutexes and unique handles (#23570 )
2026-07-21 23:40:45 +08:00
5735e10c49
ggml-openvino: Add GGML_BACKEND_DL_IMPL invocation for OpenVINO backend (#25795 )
Markus Ebner and GitHub
2026-07-21 16:43:11 +02:00
305ba519ab
CUDA: vectorize same-type get_rows with int4 copy (#25929 )
Piotr Wilkin (ilintar) and GitHub
2026-07-21 15:53:57 +02:00
76f46ad29d
hexagon: add CLAMP op (#25934 )
Todor Boinovski and GitHub
2026-07-20 16:12:09 -07:00
2beefef688
ui: Sidebar Conversations Bulk Action + Improved Settings logic/UI (#25815 )
2026-07-20 23:40:08 +02:00
91d2fc3875
llama_dsv4: write only used rows in state (#25325 )
Aman Gupta and GitHub
2026-07-20 22:43:39 +08:00
4ee6a9af71
ui: fix collapsed user bubble with markdown rendering (#25869 )
Pascal and GitHub
2026-07-20 16:28:43 +02:00
43b5e63589
UI: fix Settings/Display tool call content toggle (#25783 )
Pascal and GitHub
2026-07-20 16:28:24 +02:00
1521a9ac31
ui: enable the agentic flow when only the JS sandbox is active (#25865 )
Pascal and GitHub
2026-07-20 16:22:16 +02:00
178a6c4493
opencl: Support broadcast for Adreno MUL_MAT and honor view_offs for Adreno Q8_0 MUL_MAT for llama-server multi-stream (#25910 )
2026-07-19 22:48:57 -07:00
571d0d540d
model: rotate injected K/V cache for DFlash (#25823 )
2026-07-18 15:02:18 +02:00
4937ca83f4
llama-quant : exclude i32 ffn_gate_tid2eid routing table from quantization (#25787 )
Yash Raj Pandey and GitHub
2026-07-18 07:43:18 -04:00
86a9c79f86
opencl: load and use kernel_gemm_moe_q6_k_f32_ns from bin kernel lib (#25797 )
lhez and GitHub
2026-07-17 15:29:29 -07:00
6bdd77f13c
opencl: read/write MoE dp4a activation tiles to local memory as 128-bit (vectorized LD/ST perf opt) for Adreno GPUs (#25810 )
Hongqiang Wang and GitHub
2026-07-17 12:02:27 -07:00
86d86ed439
opencl: transpose q4_K noshuffle scales for coalesced reads (#25805 )
Hongqiang Wang and GitHub
2026-07-17 07:49:43 -07:00
7d56da7e54
sync : ggml
Georgi Gerganov
2026-07-17 16:45:30 +03:00
3727404068
ggml : bump version to 0.17.0 (ggml/1568)
Georgi Gerganov
2026-07-17 16:44:55 +03:00
5d5306bf3e
tests : initialize all tensors in test_dsv4_hc to avoid NaNs in sentinel tensors (#25822 )
2026-07-17 15:33:35 +02:00
635cdd5fcc
common : auto-download dflash- and eagle3- HF sidecars (#25811 )
Georgi Gerganov and GitHub
2026-07-17 12:15:30 +03:00
11fd0a6fb7
ggml-blas: default hadamard mul_mat to cpu routine (#25710 )
Aaron Teo and GitHub
2026-07-17 16:39:33 +08:00
788e07dc91
vulkan: Support Q2_0 (#25430 )
Jeff Bolz and GitHub
2026-07-17 07:42:59 +01:00
0bd0ec6099
sycl: fix row calculation when K_QUANTS_PER_ITERATION is 1 (#25690 )
Todd Malsbary and GitHub
2026-07-16 22:49:49 -07:00
b85833e934
opencl: add ABS op (#25115 )
Gezahegne and GitHub
2026-07-17 01:13:47 -04:00
e8f19cc0ad
opencl: loads quants as uint for q4_K and q5_K flat mv (optimization for Adreno A7x GPUs) (#25780 )
2026-07-16 13:18:21 -07:00
ac2557cb24
docs: added a note about using OpenCl with Adreno 810 (#25786 )
akleine and GitHub
2026-07-16 21:44:45 +02:00
0dc74e332e
DeepseekV4: Add fused hyper-connection ops (#25585 )
Aman Gupta and GitHub
2026-07-17 00:33:33 +08:00
b2dd28a3b6
hexagon: L2 cache handling rework (dirty bit tracking with lazy flushing) and more MUL_MAT updates (#25762 )
Max Krasnyansky and GitHub
2026-07-16 09:28:04 -07:00
f15bd60901
kleidiai: Add SME vs SME2 distinction in kernel dispatch (#25478 )
Rajendra Matcha and GitHub
2026-07-16 21:27:04 +05:30
b15ca938ad
vulkan: when using transfer queue for async copies, sync on event_wait to avoid race (#25229 )
Ruben Ortlam and GitHub
2026-07-16 15:34:24 +02:00
3278e921b1
conversion: accept BitNetForCausalLM architecture name (#25769 )
Khashayar Ghafouri and GitHub
2026-07-16 18:54:47 +05:30
2e1fd76490
TP: fix Phi3, Bert, Plamo2/3, ChatGLM (#25536 )
Johannes Gäßler and GitHub
2026-07-16 15:23:23 +02:00
86b719bf21
vendor: update BoringSSL to 0.20260713.0 (#25624 )
Alessandro de Oliveira Faria (A.K.A.CABELO) and GitHub
2026-07-16 10:17:38 -03:00
32e789fdfd
tests: actually exercise test-recurrent-state-rollback (#25758 )
Aman Gupta and GitHub
2026-07-16 21:06:12 +08:00
a8dc0e3269
server : allow text-only slot save/restore with mtmd (#25076 )
Chipmunk and GitHub
2026-07-16 21:26:44 +09:00
a55a8c5266
convert : fix dflash target tokenizer mismatch during conversion (#25733 )
Ruixiang Wang and GitHub
2026-07-16 14:19:47 +02:00
79bba02a67
CUDA: Support CUDA Virtual Devices (#25228 )
Anav Prasad and GitHub
2026-07-16 03:37:35 -07:00
3f08ef2c51
Enable CUDA graphs on volta+turing (#25749 )
Alexander Heisler and GitHub
2026-07-16 05:56:19 -04:00
8ee54c8b32
server: Ignore empty / non-existing Origin headers (#25756 )
Sebastian Dröge and GitHub
2026-07-16 12:26:51 +03:00
c7d8722922
ggml-cuda : restore prop.integrated on HIP builds (#24233 )
liminfei-amd and GitHub
2026-07-16 17:10:08 +08:00
5839ba3524
CUDA: dedup MoE gate/up activation quantization (#25441 )
2026-07-16 07:02:25 +00:00
a320cbfcb7
ci : add official website link to release notes (#25728 )
Georgi Gerganov and GitHub
2026-07-16 08:30:42 +03:00
56d6e9dde2
quant : allow using manual tensor types with --pure (#25716 )
Georgi Gerganov and GitHub
2026-07-16 08:30:20 +03:00
3dafb585f8
opencl: disable FA and MoE weights repack to work around compiler issues for Adreno 850 GPU (#25745 )
2026-07-15 20:53:14 -07:00
602f828b4d
cuda: extract Q1_0 elements via __byte_perm (#25628 )
David Friehs and GitHub
2026-07-16 05:39:17 +02:00
505b1ed15c
opencl: exclude some moe kernels on Adreno a7x (#25698 )
Hongqiang Wang and GitHub
2026-07-15 12:02:19 -07:00
32beb244f5
ui: Agentic Content UX improvements (#25450 )
Aleksander Grygier and GitHub
2026-07-15 20:31:45 +02:00
3b53219361
cuda : CUDA GGML_OP_LIGHTNING_INDEXER implementation (generic vector kernel + wmma kernel) (#25545 )
2026-07-15 19:57:52 +02:00
aff6eb6e75
tokenize : drop --stdin mutual-exclusion check (#25672 )
Adrien Gallouët and GitHub
2026-07-15 18:41:51 +02:00
c3d47e696b
opencl: fix two issues on flash attention for Adreno a7x (#25697 )
Hongqiang Wang and GitHub
2026-07-15 09:08:40 -07:00
f6f12e43fa
CUDA: tighter MMQ src1 buffer size for native fp4 (#25613 )
leonardHONG and GitHub
2026-07-15 23:21:22 +08:00
956973c764
Fix crash with draft-simple (#25720 )
Gaurav Garg and GitHub
2026-07-15 19:51:34 +05:30
a582222290
server: fix read_file append_loc space breaking edit_file match (#25705 )
Pascal and GitHub
2026-07-15 13:46:46 +02:00
a05df0a81a
ui: fix thinking menu never appearing in single-model mode (#25637 )
Pascal and GitHub
2026-07-15 13:39:21 +02:00
a3e5b96ac5
cuda : relax tensor contiguity requirements for quantized concat (#25678 )
2026-07-15 13:36:32 +02:00
c81029373d
ci : add HF_TOKEN to self-hosted workflows (#25706 )
2026-07-15 14:34:53 +03:00
b3c9d1b846
metal: fuse snake activation (mul, sin, sqr, mul, add) (#25459 )
Pascal and GitHub
2026-07-15 12:53:31 +02:00
f955e394bf
ggml: add f16 out_prod support for CPU and out_prod op for Vulkan (#23997 )
Michael Lamothe and GitHub
2026-07-15 18:46:56 +10:00
33a75f41c3
DeepseekV4: reduce graph splits (#25702 )
Aman Gupta and GitHub
2026-07-15 15:47:18 +08:00
d3fba0c79d
sycl : fix get_rows Q2_K, Q4_K, Q5_K (#25656 )
Neo Zhang and GitHub
2026-07-15 15:32:28 +08:00
ae9291e16b
sycl : support kernel type fp16 for conv2d_dw (#25653 )
Neo Zhang and GitHub
2026-07-15 15:31:10 +08:00
22b208b1ca
sycl : implement xielu op (#25550 )
Andrew Smith and GitHub
2026-07-15 00:29:12 -07:00
0e148a573f
sycl: Increase minimum buffer size for USM system allocations (#25525 )
Francois Dugast and GitHub
2026-07-15 09:28:24 +02:00
32b741c336
[SYCL] Flash Attention with XMX engine via oneDNN (#25222 )
2026-07-15 03:26:53 -04:00
12127defda
opencl: do not use clCreateBufferWithProperties when targeting CL 2.x (#25673 )
Hongqiang Wang and GitHub
2026-07-14 19:53:56 -07:00
00fa7cb284
opencl: handle OOB write in noshuffle GEMV kernels (odd ne01) (#25640 )
Hongqiang Wang and GitHub
2026-07-14 13:46:54 -07:00
a4ce2595c5
opencl: avoid the vec path in GEMV for unaligned row stride (#25671 )
Hongqiang Wang and GitHub
2026-07-14 12:27:56 -07:00
c71854292f
hexagon: fix hmx-queue signal enum-narrowing problem (#25677 )
Chyan and GitHub
2026-07-15 03:27:09 +08:00
bf2c86ddc0
server : refactor prompt cache state ownership (#25649 )
Georgi Gerganov and GitHub
2026-07-14 18:25:52 +03:00
6e52db5b72
server: add --cors-* options (#25655 )
Xuan-Son Nguyen and GitHub
2026-07-14 17:23:44 +02:00
236ab574e0
ui: Fix spacing in tool-call request (#25634 )
Bill Sideris and GitHub
2026-07-14 18:23:11 +03:00
dfba90db63
webui: parse effective-parameter sizes (E2B, E4B) as params (#25529 )
Emanuil Rusev and GitHub
2026-07-14 18:12:22 +03:00
00e79f6fb1
opencl: fix a dp4a bug for devices where cl_khr_integer_dot_product is unavailable (#25639 )
2026-07-14 08:08:13 -07:00
17a05e451f
ui: fix mcp panel for toggle + timeout + proxy + ON/OFF state (#25631 )
Pascal and GitHub
2026-07-14 16:50:44 +02:00
7f575c39d6
DeepseekV4: fix seq_rm (#25588 )
Aman Gupta and GitHub
2026-07-14 21:45:36 +08:00
7cbd61002d
vulkan/cpu: Support f16 as SET_ROWS src. (#25432 )
Jeff Bolz and GitHub
2026-07-14 08:26:55 -05:00
8ff8c4299d
tokenize : align usage by using common args (#25516 )
Adrien Gallouët and GitHub
2026-07-14 15:20:53 +02:00
a7312ae94f
ggml : add a set of functions for checking contiguity of inner tensor dimensions (#25650 )
2026-07-14 14:37:52 +02:00
657e01125a
tests: export-graph-ops: exit gracefully when called w/o arguments (#25619 )
Christian Kastner and GitHub
2026-07-14 12:15:41 +02:00
47a39665e7
ggml: uniformize im2col dst_type for all conv ops (#23660 )
2026-07-14 12:13:13 +02:00
47c786924a
kleidiai : add SME2 f32 kernel (#24414 )
Charles Xu and GitHub
2026-07-14 12:12:18 +02:00
c9330ed0cf
ui: add reasoning effort control to mobile add sheet (#25539 )
Pascal and GitHub
2026-07-14 12:05:40 +02:00
cb489bc0fb
convert_hf_to_gguf: support split MTP export for HY V3 (#25641 )
Thiago Padilha and GitHub
2026-07-14 06:43:15 -03:00
ec0dbef816
arg: Flush log before exiting after usage() (#25504 )
Christian Kastner and GitHub
2026-07-14 11:03:22 +02:00
c1063ac9d7
sycl: set fattn_vec_nthreads to 256 for Battlemage (#25205 )
Titaniumtown and GitHub
2026-07-14 05:00:00 -04:00
14d3ba45f3
metal : add Q2_0 support (#25419 )
Pasha Khosravi and GitHub
2026-07-13 21:52:00 -07:00
2969d6d15d
model: add Hy3 (hy_v3) support with MTP speculative decoding (#25395 )
2026-07-14 10:31:04 +12:00
6eddde06a4
CUDA: refactor MMQ kernel configuration (#24127 )
Johannes Gäßler and GitHub
2026-07-13 18:37:57 +02:00
e920c523e3
vulkan: Use native e2m1 and e4m3 conversions for mxfp4/nvfp4 (#25338 )
Jeff Bolz and GitHub
2026-07-13 08:44:17 -05:00
259ae1df8b
spec: add Minimax2 eagle3 support
2026-07-13 06:22:37 -07:00
4193ea697f
readme : add link to maintainer PRs (#25621 )
Georgi Gerganov and GitHub
2026-07-13 16:07:58 +03:00
f4253ef965
tests: Harmonize header use (#25616 )
Christian Kastner and GitHub
2026-07-13 14:36:51 +02:00
ad8d821991
gguf : add tensor shape accessor (#24405 )
QuintinShaw and GitHub
2026-07-13 18:55:15 +08:00
91c631b21d
chat : fix reasoning leak with force-opened bare <think> templates (#24674 )
2026-07-13 02:45:10 -05:00