Commit Graph
5429 Commits
Author SHA1 Message Date
Masashi Yoshimura 51bfbd6170 webgpu: add f16 support to fill/set_rows (llama/29897) 2026-10-06 10:38:19 +03:00
f2aa80ad26 ggml-openvino: update to 2026.4.1, optimize performance, expand ops, improve device listing. (llama/29852)
* ggml-openvino : Qwen3.5 MoE perf (llama/312)

Squash of ravi9/llama.cpp#312:

- ggml-openvino: add detailed inference profiling (Yu, Zijun)
- ggml-openvino: use remote output tensors by default (Yu, Zijun)
- ggml-openvino: optimize single-sequence recurrent state (Yu, Zijun)
- opt1: remove recurrent reset for single sequence, opt2: direct gdn outputs (break parallel sequence) (Yu, Zijun)
- fix parallel sequences (Yu, Zijun)
- ggml-openvino: simplify graph cache key (ynimmaga)
- enable stateful for qwen35 single sequence (Yu, Zijun)
- Fix after rebasing (Yu, Zijun)
- Add k-requant option q4_asym64 (Yu, Zijun)
- Fix qwen35 llama-bench -p 0 (Yu, Zijun)
- Simplify RESHAPE translation (Yu, Zijun)
- openvino: fuse MoE routing (Yu, Zijun)
- openvino: fuse GDN qk normalization (Yu, Zijun)
- openvino: enable GPU MoE fusion by default (Yu, Zijun)
- ggml-openvino: add cache_only mode to import cached compiled model on disk directly (Yu, Zijun)
- openvino : report the device allocation limit to ggml (Łukasz Ślusarczyk)
- Fix windows build (Yu, Zijun)

Co-authored-by: ynimmaga <ynimmaga@users.noreply.github.com>
Co-authored-by: Łukasz Ślusarczyk <lukasz.slusarczyk@intel.com>

* ggml-openvino: Update doc of compiled model cache

* openvino: implement PRD-compliant device enumeration and memory reporting

* openvino: fix multi-device listing issues from review

- Only the device selected by GGML_OPENVINO_DEVICE reports as GPU; the
  other OpenVINO devices report as IGPU so llama.cpp does not offload to
  them. Initializing a non-selected device logs a warning.
- Name devices OPENVINO<i> again and show the OpenVINO id in the
  description. Raw "CPU" names shadowed the ggml CPU backend.
- Support GPU.N: create the OpenCL queue on OpenVINO's own context for
  the selected device, and replace "GPU"/"NPU" string comparisons with
  ggml_openvino_is_gpu()/ggml_openvino_is_npu().
- An unavailable GGML_OPENVINO_DEVICE is now an error that lists the
  available devices, instead of silently falling back to CPU.
- Memory: cap iGPU/NPU free memory at system available memory, fall back
  to system memory instead of 0/0 when the plugin lacks memory
  properties, and ignore host USM allocations in GPU usage.
- Initialize the device config once under a lock, even if OpenCL setup
  fails.
- Fix supports_op return type for non-selected devices (build error).

* openvino : take USM entry points from the selected device platform

clGetExtensionFunctionAddressForPlatform was called on the first platform
returned by clGetPlatformIDs. The address it returns is only valid for the
platform it was queried on, and the first platform is not always the one that
holds the device OpenVINO selected.

On a host whose first platform comes from another vendor the lookup returns
null, and then every read, write and memset on a GPU buffer fails with
"clEnqueueMemcpyINTEL not available".

Look both entry points up in init(), on the platform of the device OpenVINO
picked, and keep them in the device config next to the command queue.

Assisted-by: Claude Opus 5

* openvino: fuse MoE experts for models with a fused gate_up weight

FuseMoeCompressed only matches models whose gate and up projections are
separate GatherMatmul ops. gemma-4 packs both into one expert weight and
splits the result after the GEMM, so its MoE block stayed unfused and ran
the expert GEMMs as per-token GEMVs.

Add FuseMoeCompressedFusedGateUp, which matches that shape
(one GatherMatmul -> Slice/Slice -> Gelu(ERF) -> Multiply) and folds it into
the same MOECompressed op, using GEMM3_SWIGLU with GEGLU_ERF. The fused
weight, scale and zero point are split into gate/up halves by copying raw
bytes, since a graph Slice would be rewritten to StridedSlice and constant
folded, whose reference evaluator crashes on sub-byte types.

gemma-4 also applies a per-expert output scale to the down projection before
the router weights. MOECompressed takes only one per-expert weight, so that
scale is folded into the routing weights, which is exact.

The op reads the zero point straight off a weight port and needs an integer
Constant there, so the matcher requires one and leaves natively quantized
experts (exact f16 zp) to the unfused path.

gemma-4-26B-A4B on Arc B390, GGML_OPENVINO_REQUANT_KQUANT=q4_asym64_all,
llama-bench -p 512 -n 128 -r 2, against a GGML_OPENVINO_MOE_OP=0 baseline:
pp512 66.16 -> 1608.73 t/s, tg128 25.94 -> 26.46 t/s. Perplexity over 12
chunks is unchanged (1451.3 +/- 177.9 unfused vs 1427.6 +/- 175.1 fused).

No effect without that requant option, on models with separate gate/up
weights, or on CPU. test-backend-ops -b OPENVINO0 is unchanged by this
commit: two MUL_MAT_ID m_v cases fail, the same two on the unmodified base.

* openvino: fix rank-3 axis handling so MoE works under stateful execution

Stateful execution drops the leading size-1 batch dim, so OV tensors are rank
3 while GgmlOvDecoder::get_shape/get_stride still report GGML_MAX_DIMS=4
reversed entries. Several MoE ops derive OV axis indices straight from that
metadata, so they picked the wrong axis. A MoE model with
GGML_OPENVINO_STATEFUL_EXECUTION=1 aborts while building the graph:

  Check 'is_axis_valid(axis, r)' failed at src/core/src/validation_util.cpp:336
  While validating node 'opset11::TopK ... _ffn_moe_probs ...'
  Axis 3 out of the tensor rank range [-3, 2].

Fix idiom throughout: take the axis from the real OV rank, or shift a
metadata-derived axis down by metadata_rank - actual_rank.

  argsort.cpp    the router top-k axis is 2 on rank 3, not 3. This is the
                 abort quoted above.
  add.cpp        the MoE expert-sum bypass collapses the 8-ADD chain into one
                 ReduceSum on hardcoded axis 2, which on rank 3 reduces n_embd
                 instead of the expert axis. Now rank-2, with the following
                 Unsqueeze at rank-3.
  get_rows.cpp   squeezing a hardcoded {0,1} also strips the batch dim
                 whenever it is 1, which is every decode step. Squeeze down to
                 the trailing two dims instead.
  mul_mat_id.cpp pick the reshape dims by actual rank, and skip the trailing
                 Unsqueeze that re-adds the batch dim.
  view.cpp       the expert-plane slice had the Slice axis, dst_ov_axis, the
                 ShapeOf+Gather index and the Reshape target all rank-4.
  utils.cpp      process_view_input_new's "translate_view already resolved
                 this VIEW, skip re-slicing" shortcut required equal ranks. 4
                 vs 3 never matched, so every resolved expert plane got
                 re-sliced. Now compares the common trailing dims. Same axis
                 shift for the Slice in the view-chain walker.

Stateless is unchanged by construction: every edit is gated on the actual
rank, so axis_shift == 0 reproduces the previous code exactly. Checked on
OV-CPU by diffing greedy output against the unmodified base for dense
gemma-4-E2B, granite-1b-a400m and gemma-4-26B-A4B; all identical.

granite-1b-a400m on OV-CPU aborts with the error above before this change;
after it, it generates and is byte-identical to stateless. Dense gemma-4-E2B
is identical stateless vs stateful both before and after. test-backend-ops
-b OPENVINO0 is unchanged: two pre-existing MUL_MAT_ID m_v cases fail, the
same two on the unmodified base.

gemma-4-26B-A4B is a poor correctness vehicle here. On OV it already drifts
into degenerate repetition a few tokens in, in stateless as much as stateful,
and the two modes diverge somewhere inside that degenerate region instead of
matching token for token. Each mode is self-reproducible across runs.

Known limitation: FuseMoeCompressedFusedGateUp does not match the rank-3
graph, so a MoE model run with GGML_OPENVINO_STATEFUL_EXECUTION=1 loses the
prefill fusion while gaining decode. gemma-4-26B-A4B on Arc B390,
GGML_OPENVINO_REQUANT_KQUANT=q4_asym64_all, llama-bench -p 512 -n 128 -r 2:

  unfused (GGML_OPENVINO_MOE_OP=0)  pp512   66.16   tg128  25.94
  fused, stateless (default)        pp512 1608.73   tg128  26.46
  fused, stateful                   pp512   66.18   tg128  29.91

Stateful is opt-in and off by default, and MoE did not run there at all
before this, so nothing that previously worked regresses. Making the pass
match rank 3 is the follow-up.

* OpenVINO Backend: Upgrade graph cache to use node_idx, src_idx, node type

* ggml-openvino : enable more comprehensive conv fusion

* enable conv ops

* Reject kernel size 0 and support IM2COL_3D

* openvino : abort when the GPU remote context cannot be created

init() logged the error and returned, which left the device name a GPU but
remote_context empty. The remote buffer and tensor paths assert only on the
device being a GPU and then dereference that empty optional.

Those paths have no host fallback, and a device that OpenVINO listed should
have a working OpenCL context, so stop instead of continuing. An OpenCL stack
that is broken as a whole is still caught earlier by the device availability
check, which falls back to CPU.

Assisted-by: Claude Opus 5

* openvino : fix build warnings

The single-argument form of the OpenVINO RTTI macros is the intended one, but
their selector macro leaves __VA_ARGS__ empty, which -Wpedantic reports on
every pass and op header. Turn that warning off for this backend only, the
way ggml-cuda and ggml-sycl already do for their own third-party warnings.

Also drop a break and a dead assignment around a GGML_ABORT, which is noreturn.

Assisted-by: Claude Opus 5

* OpenVINO Backend: Support common MTMD ops

* ggml-openvino: give a reshaping view its own ov::Tensor

* ggml-openvino : compute HARDSIGMOID and EXPM1 in f32

HARDSIGMOID used a 1/6 constant in the input type, which is not exact
in bf16, and EXPM1 lost precision for small inputs in f16. Both now
compute in f32 and convert back, except on NPU where the f32 path
gives wrong results.

Fixes the HARDSIGMOID/EXPM1 test-backend-ops failures on GPU.

* ggml-openvino : update device selection and --list-devices

Show the selecting GGML_OPENVINO_DEVICE value and active device in
--list-devices, startup logs, and backend tests.

Clarify OpenVINO selection uses GGML_OPENVINO_DEVICE, not -dev.

* openvino : remove unreachable OpenCL queue checks

A remote buffer exists only on a GPU device, and init() aborts there if the
queue cannot be created, so the queue is never null at these call sites.

Assisted-by: Claude Opus 5

* openvino : update OpenVINO to 2026.4.1 and GPU drivers to 26.35.39758.10

* docs : update OpenVINO validated models and GPU driver version

* ggml-openvino : skip empty views when giving a reshaping view its own tensor

A zero-size view can sit at the end of a GPU USM buffer (Qwen3.5 recurrent cache). Wrapping it as a remote tensor throws "shared USM buffer has smaller size (0)".

Assisted-by: Claude

* ggml-openvino : rebind the cached decoder when llama passes a different graph

llama keeps separate graphs for batches with and without outputs. llama-server splits the prompt into chunks for context checkpoints, so a cached decoder could be reused with a graph built in other memory and bind the previous chunk's input tensors. SWA and recurrent models then lost most of the prompt in llama-cli and llama-server.

Assisted-by: Claude

* docs : update OpenVINO validated models

Smoke test on Lunar Lake (32 GB) with the two fixes above. Re-add the Qwen3.5 and gemma models.

Assisted-by: Claude

---------

Co-authored-by: Yu, Zijun <zijun.yu@intel.com>
Co-authored-by: ynimmaga <ynimmaga@users.noreply.github.com>
Co-authored-by: Łukasz Ślusarczyk <lukasz.slusarczyk@intel.com>
Co-authored-by: haarika-madaka <haarika.madaka@intel.com>
Co-authored-by: Mustafa Cavus <mustafa.cavus@intel.com>
Co-authored-by: Mostafa Faheem <mostafaaafaheem@gmail.com>
2026-10-06 10:38:18 +03:00
PascalandRuben Ortlam 7f4a67a504 qwen4exp : halve the indexer score memory (llama/29825)
* qwen4exp : halve the indexer score memory

The indexer scored all heads in one product and rectified a copy of it,
so two [n_pool, n_idx_h, n_tokens] f32 tensors were live at once, the
largest buffers of the graph at long context. Each head now gets its
own product, rectified and summed in place into one [n_pool, n_tokens]
score.

* qwen4exp: let the allocator reuse the indexer score buffers

Address review from CISC: use plain ggml_add and ggml_relu in the
indexer head loop. The graph allocator already runs them in place when
their source has no other consumer, so the _inplace variants are not
needed. The compute buffer and the speed are unchanged.

* cuda: support 4 heads in the lightning indexer

Dispatch 4 heads to the vector kernel, too few for a wmma tile, and
accept them in supports_op. test-backend-ops covers 4 heads.

* metal: take the lightning indexer head count as a function constant

The kernel reads the head count from a function constant and zero fills
the last head tile, so any head count runs and 64 heads is unchanged.

* qwen4exp: compute the indexer score with the lightning indexer

Address review from am17an: the unweighted sum of the rectified head
scores scaled by 1/sqrt(head_dim) is the lightning indexer with every
head weight set to that scale, so the indexer calls
ggml_lightning_indexer on the pooled keys with an f16 pool mask. The
keys are read once for all heads and no per head score is
materialized.

* vulkan: tile the lightning indexer over keys and tokens

A workgroup scores 64 keys against 8 tokens: the keys are staged once
in shared memory, the queries one head at a time, and each invocation
owns one key for two tokens, so no dot product needs a cross invocation
reduction. The subgroup variant and the flat dispatch are gone, the grid
is keys x tokens x streams.

* vectorize vulkan loads and use fp16 dot product

---------

Co-authored-by: Ruben Ortlam <rortlam@redhat.com>
2026-10-06 10:38:18 +03:00
Aman Gupta cf5d9e3014 CUDA: fuse shared experts into MMVQ (llama/29184)
* CUDA: fuse shared experts into MMVQ

* check if buffer is null

* move stride_col_dst to fusion args
2026-10-06 10:38:18 +03:00
Yash Raj Pandey d7682b5814 ggml-cuda : fix cpy transposed path corrupting non-contiguous dst (llama/27663) 2026-10-06 10:38:18 +03:00
Yash Raj Pandey 6b704713cc ggml-quants : avoid invalid rounding in qkx3 scale search (llama/29817)
* ggml-quants : avoid invalid rounding in qkx3 scale search

The imatrix scale search can produce an infinite, NaN, or otherwise out-of-range value when the fitted minimum collapses to the maximum or makes the range extremely small. That value is then passed to nearest_int and can trip its assertion in Debug builds.

Clamp the quantization level to [0, nmax] before rounding so valid in-range values behave the same as before while invalid scale-search results no longer reach nearest_int.

Add regression coverage for degenerate imatrix groups across q2_K, q4_K, q5_K, q4_1, and q5_1.

Fixes #29804.

Assisted-by: Claude Opus 5.5

* tests: print degenerate imatrix quant types
2026-10-06 10:38:18 +03:00
Yash Raj PandeyandGeorgi Gerganov 0b35d1f30b ggml-cpu : fix soft_max_back wrong output when dst aliases src1 (llama/27096)
* ggml-cpu : fix soft_max_back wrong output when dst aliases src1

GGML_OP_SOFT_MAX_BACK is listed in ggml_op_can_inplace, so the graph
allocator may assign dst to alias either src0 (dy) or src1 (y).

The result was built in several steps:

    ggml_vec_cpy_f32  (nc, dx, dy);
    ggml_vec_acc1_f32 (nc, dx, -dot_y_dy);
    ggml_vec_mul_f32  (nc, dx, dx, y);
    ggml_vec_scale_f32(nc, dx, scale);

When dst aliases src1, the first step overwrites y and the third step
then reads the overwritten values, so the output is silently wrong.
Aliasing dst with src0 is unaffected. The CUDA kernel completes its
reduction before writing and is already safe.

Replace the sequence with a single fused loop that reads both sources
before writing, which is correct under either aliasing.

Add a regression test that marks dy as a graph output so the allocator
is forced to alias dst with y, asserts that the alias actually
happened, and compares against values computed on the host.

* cont : remove comment

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
2026-10-06 10:38:18 +03:00
Ethan Guo d0dc436345 metal : add tensor API flash attention kernel for F16 KV (llama/29570)
* metal : add tensor API flash attention kernel for F16 KV

* cont : add tensor FA kernels for DK=DV=512 and DK=576, DV=512

* cont : support attention sinks, ALiBi and logit softcap in the tensor FA kernel

* cont : add tensor FA kernel for DK=192, DV=128
2026-10-06 10:38:17 +03:00
lhez cca8f72b46 opencl: use sigmoid f16 for bf16 (llama/29787) 2026-10-06 10:38:17 +03:00
cwriterandcwriter cd1bee5203 SYCL: Q8_0 DMMV ESIMD and MMVQ wide load (llama/29186)
* Adding wide-load mmvq for Q8_0 and esimd dmmv for q8_0

Assisted-by: Codex

* remove guard for q8_0

* remove docs

* Simplify by committing to clean code without fallback

* Add feature flag as requested

Assisted-by: Claude Opus 5

---------

Co-authored-by: cwriter <cwriter@localhost>
2026-10-06 10:38:17 +03:00
Jiwoong Song 95a05ed4da vulkan: disable large matmul tile on Samsung GPUs with 32KB shared memory (llama/28531)
Assisted-by: Claude Opus
2026-10-06 10:38:17 +03:00
Titaniumtown c4051a33e6 sycl: large register file for D=512 FA vec kernels (llama/29062)
* sycl: large register file for D=512 FA vec kernels

* tests: add 512-wide FA heads to the perf sweep
2026-10-06 10:38:17 +03:00
Łukasz Ślusarczyk ae92605fca sycl : do not use slow oneDNN reference matmul and fattn (llama/28985)
* sycl : do not use slow oneDNN reference matmul and fattn

* sycl : probe oneDNN matmul once at device init

Assisted-by: Claude Opus 5
2026-10-06 10:38:16 +03:00
Georgi Gerganov 9c968f7e78 qwen4exp : optimize mask constructions (llama/29824)
* qwen4exp : optimize mask constructions

* cont : apply the same change for GLM5-next
2026-10-06 10:38:16 +03:00
Georgi Gerganov 4435763c17 ggml : add alloc_buffer_n to buffer type interface (llama/23671)
* ggml : add `alloc_buffer_n` to buffer type interface

Add alloc_buffer_n method to ggml_backend_buffer_type_i
interface, with a public API ggml_backend_buft_alloc_buffer_n.

- Default implementation in ggml-backend.cpp handles multi-buffer
  splitting and tensor allocation via ggml_tallocr
- Meta buffer type provides custom implementation that creates
  per-device sub-contexts and delegates to simple buffer types
- ggml_backend_alloc_ctx_tensors_from_buft now collects tensors
  into a list and delegates to the new API
- Remove temporary ggml_backend_meta_alloc_ctx_tensors_from_buft
- Add NULL alloc_buffer_n to all existing buffer type
  interfaces (cpu, metal, openvino, hexagon, webgpu, zdnn, virtgpu, repack)

Assisted-by: llama.cpp:local pi

* cont : fix `cur_buf_size` init after flushing a buffer

* ggml : add TODO tag for shared buffer split logic

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* tests : add alloc_buffer_n coverage

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* cont : fix compile warnings

* tests : add descriptions for alloc_buffer_n tests

Assisted-by: pi:llama.cpp/Qwen3.8-27B

* ggml : address review comments on alloc_buffer_n

- restore GGML_LOG_ERROR on buffer alloc / tensor init failure in the
  default impl (name the failing tensor)
- check the malloc result and drop the _impl indirection in
  ggml_backend_alloc_ctx_tensors_from_buft
- remove comments that restate the code
- fix the TAG_ALLOC_SHARED_BUFFER_SPLIT typo

Assisted-by: pi:llama.cpp/Qwen3.8-27B

* ggml : add get_alloc_size_n to buffer type interface

- Add ggml_backend_buft_get_alloc_size_n public API
- Add optional get_alloc_size_n callback to ggml_backend_buffer_type_i
- Share tensor->buffer planning between alloc_buffer_n default and get_alloc_size_n default
- Replace unchecked realloc with std::vector in alloc_buffer_n default
- Make ggml_backend_alloc_ctx_tensors_from_buft_size use the new API
- Add test-alloc coverage for get_alloc_size_n

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* cont : report malloc failure
2026-10-06 10:38:16 +03:00
Ruben Ortlam 8291ab849c vulkan: add logging to pipeline compile issues (llama/29794) 2026-10-06 10:38:16 +03:00
kurquhar eaadab3d53 hexagon: install rebuilt HTP skels (llama/29828)
* hexagon: install rebuilt HTP skels

Assisted-by: OpenCode

* hexagon: fix HTP skel catalog dependencies

Assisted-by: OpenCode
2026-10-06 10:38:16 +03:00
Jhen-Jie HongandMax Krasnyansky 0295ef6846 hexagon: add q2_k and q3_k quant type support (llama/29717)
* hexagon: add q2_k and q3_k quant type support

* hex-qk: consistent allocation of src1_row_size

---------

Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
2026-10-06 10:38:15 +03:00
Johannes Gäßler 396f68d33e CUDA: fix 2 broken Volta FA cases (llama/29803) 2026-10-06 10:38:15 +03:00
Johannes Gäßler 3a3598cc24 llama: refer to segment documentation [no ci] (llama/29074) 2026-10-06 10:38:15 +03:00
Yiwei ShaoandMax Krasnyansky 15229ac95d hexagon: shared strided DMA copy for CPY and CONCAT, any-dim CONCAT via DMA (llama/29685)
* hexagon: shared strided DMA copy for CPY and CONCAT, any-dim CONCAT via DMA

* hex-cpy: various fixes on top of the concat optimizations

Removed CONCAT_DMA_MIN_ROW logic, it was broken with 64-bit DMA.
While it's kinda silly to use DMA for tiny stuff if that tensor gets mapped to an extended buffer the only way to read it is DMA.

Added missing dma_queue_flush() calls.

Added additional guards for conditions we don't support.

---------

Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
2026-10-06 10:38:15 +03:00
aa538005fc cuda : route sm70 to the Turing MMVQ nwarps table (llama/29753)
* cuda : route sm70 to the Turing MMVQ nwarps table

Volta (sm_70) has no MMVQ parameter table of its own and falls through
to GENERIC, which launches K-quant batch-1 decode (ncols_dst == 1) at
nwarps=4. sm_70 shares TURING's tuning: the K-quant vec_dot prefers
nwarps=2 there. Route sm_70 to the existing MMVQ_PARAMETERS_TURING
table in both the device and the host table selector.

Measured on one Tesla V100 32GB PCIe (PG500-216, driver 580.178.04,
CUDA 12.0.140) with Qwen3.8-27B Q4_K_M, tg128, interleaved A/B in 6
ABBA blocks with paired per-block deltas: +1.091 t/s = +3.17 %
(t = +49.0, all six per-block deltas positive); perplexity
bit-identical (6.3697 +/- 0.04066 both builds, wiki.test.raw). The
patched build's K-quant mul_mat_vec_q kernels launch at nwarps=2
(cubin EIATTR_MAX_THREADS) while Q4_0/Q8_0 stay at nwarps=4, and the
same measurement on the September master base gave +3.84 % (t = 85).

The tuning originates from the V100-focused fork anyei/llamacpp-v100
(MIT), commit b912d1b1e, which carries a dedicated
MMVQ_PARAMETERS_VOLTA table; a cubin-level comparison confirmed that
routing sm_70 to the existing TURING table is equivalent for the
K-quant batch-1 path this change affects, so this is the minimal
2-line form. https://github.com/anyei/llamacpp-v100/commit/b912d1b1e

Original-patch-by: anyei <angelyoelroblesmercedes@gmail.com>

* Update ggml/src/ggml-cuda/mmvq.cu

---------

Co-authored-by: tkittich <tkittich@gmail.com>
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
2026-10-06 10:38:15 +03:00
Mike van LammerenandNiklas Wenzel 81ca4f81fa metal : release temporary private transfer buffers (llama/29777)
* metal : release temporary private transfer buffers

Assisted-by: OpenAI Codex

* metal : fix order and formatting

---------

Co-authored-by: Niklas Wenzel <dev@nikwen.de>
2026-10-06 10:38:15 +03:00
Masashi Yoshimura 88948bdb83 webgpu: add bfloat16 support for MUL_MAT/MUL_MAT_ID/GET_ROWS- #29358 (llama/29358) 2026-10-06 10:38:14 +03:00
ynankaniandJohannes Gäßler 10872be7f8 CUDA: Handle compute type for NVFP4 on cublass path (llama/29173)
* CUDA: Handle compute type for NVFP4 on cublass path

Signed-off-by: ynankani <ynankani@nvidia.com>

* Use BF16 compute type for quantized models if HW allows

Signed-off-by: ynankani <ynankani@nvidia.com>

* Set acc prec to bf16 for nvfp4 as it needs atleast bf16 range

Signed-off-by: ynankani <ynankani@nvidia.com>

* Update ggml/src/ggml-cuda/ggml-cuda.cu

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>

* preserve op_params for per-expert matmul

Signed-off-by: ynankani <ynankani@nvidia.com>

---------

Signed-off-by: ynankani <ynankani@nvidia.com>
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
2026-10-06 10:38:14 +03:00
Oliver Simons 3b68f9014c CUDA: Make CCCL configurable + pin it to 3.4.3 for CI jobs (llama/29792)
Pinning to >= 3.4.3 is required to enable DeviceTopK, which was affected by
a race condition https://github.com/NVIDIA/cccl/pull/10627.

We will relax this for future CTK versions which will bundle CCCL >
3.4.X (CTK 13.5 will bundle CCCL 3.5.0 for example)
2026-10-06 10:38:14 +03:00
Georgi Gerganov 56500f4f74 meta: clear inactive AllReduce shards with FILL, not SCALE (llama/29793) 2026-10-06 10:38:14 +03:00
uvos 817294a93e HIP: avoid treating CDNA as dgx spark for gqa_ratio 20 in fattn_mma dqk 576 (llama/29572) 2026-10-06 10:38:14 +03:00
Max Krasnyansky b4a085a024 hex-workqueue: fix race condition in seqn getting out of sync with idx_read/write (llama/29785) 2026-10-06 10:38:14 +03:00
Pradeep Rao 0ecc317e30 BLAS : Document AOCL-BLAS build and label the device AOCL-BLAS (llama/29640)
* BLAS : Document AOCL-BLAS build and label the device AOCL-BLAS

* AOCL-Blas : Add an AOCL-BLAS Quick Start and drop the fixed version path

* AOCL-BLAS doc : Note on ZenDNN
2026-10-06 10:38:13 +03:00
lhez 67ad86d020 opencl: mark vec subgroup bcast as supproted for Adreno E17 compiler (llama/29698) 2026-10-06 10:38:13 +03:00
Georgi Gerganov 61ea5027d5 metal : use bf16 math for mxfp4 mul-mat (llama/29770) 2026-10-06 10:38:13 +03:00
Masashi Yoshimura d371e376cb webgpu: fix SSM_SCAN binding aliasing (llama/29750) 2026-10-06 10:38:13 +03:00
Adrien Gallouët c1c139084c ggml-opencl : replace alloca() with std::vector (llama/29765)
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
2026-10-06 10:38:13 +03:00
R0CKSTAR bffb6e4b0f cuda: guard the iq4_nl dequantize row kernel against short rows (llama/29683)
dequantize_block_iq4_nl writes QK_K values per block, but a row can be shorter than that (an IQ4_NL row is only guaranteed to be a multiple of QK4_NL). Threads whose 32-value sub-block starts at or past k currently read and write past the end of the row. Skip those sub-blocks; for rows that are a multiple of QK_K the check never fires.
2026-10-06 10:38:13 +03:00
Ehsan BateniandMax Krasnyansky 6773ef8739 Hexagon: optimize ALLREDUCE with support for safe scatter mode (llama/29757)
* hex-allreduce: add support for safe scatter mode

* hex-allreduce: pare down excessive comments

* hex-allreduce: re-write to remove register spills

---------

Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
2026-10-06 10:38:12 +03:00
Daniel Kuts 02ea0c2ea8 ggml/gguf : fix integer overflow (llama/29384)
* ggml: fix integer overflow guard for zero-element tensors

* ggml: validate number of elements in tensor to prevent integer overflow

* ggml: fix error print
2026-10-06 10:38:12 +03:00
Adrien Gallouët 067bb06bd8 ggml-et : remove useless alloca() (llama/29663)
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
2026-10-06 10:38:12 +03:00
Aman Gupta fd66b6ecce ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA) (llama/29675)
* ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA)

* ggml-cpu : use per-op _bf16 functions for BF16 unary and GLU ops

Assisted-by: Claude Opus 5.5

* CUDA: use ggml_cuda_cast in binbcast and unary kernels to fix the HIP bf16 build

* ggml-openvino : reject BF16 SCALE and mixed-type BF16 ADD/MUL/SUB
2026-10-06 10:38:12 +03:00
Pascal c1b3fb1094 cpu: accept BF16 in src1 of mul_mat (llama/28937)
* cpu: accept BF16 in src1 of mul_mat

ggml_conv_1d_dw builds its im2col in F32 when the kernel is BF16, then
calls ggml_mul_mat(im2col, kernel), which puts F32 in src0 and BF16 in
src1. The CPU backend refused that combination, so it was reported as
unsupported on every backend and never compared against anything.

Widen BF16 into the F32 work buffer, next to the existing packing of F32
into vec_dot_type. This is the arithmetic the Metal mat vec kernel
already uses, both operands promoted to float and accumulated in float,
so the two agree exactly rather than approximately.

Cover it with a conv_1d_dw test over F32, F16 and BF16 kernels, plus
three mul_mat cases with BF16 in src1.

* vulkan: reject BF16 in src1 of mul_mat unless src0 is BF16

supports_op only checked the src1 type for non contiguous tensors, so
a contiguous BF16 src1 was accepted and the pipeline lookup asserted.
The only BF16 src1 path is the BF16 x BF16 multiply, every other src0
type now reports the op as unsupported and the scheduler keeps it on
the CPU.

The BF16 kernel case of the conv_1d_dw test needs the f32 x bf16
mat vec variants of the Metal backend, which land separately.
2026-10-06 10:38:12 +03:00
Pascal d0af734dfd openvino: serve GET_ROWS on a weight view from the base Constant (llama/28381)
* openvino: serve GET_ROWS on a weight view from the base Constant

Resolve view_src when collecting weight Constants so a view over a
quantized weight no longer becomes a dynamic typed Parameter, and fold
the row offset of the view into the gather indices instead of slicing
the dequantization subgraph.

* openvino: lift the quantized GET_ROWS view rejection

The supports_op rejection of a quantized src0 view with a nonzero
offset keeps the vs0 GET_ROWS cases of #28253 away from OpenVINO.
The weight view now resolves to the base Constant with the row offset
folded into the gather indices, so the rejection goes away.
2026-10-06 10:38:11 +03:00
R0CKSTAR 1cbf7e79f0 musa : define __CUDA_ARCH__ for device passes (llama/29508)
The MUSA vendor header never defined __CUDA_ARCH__, so every architecture
test in the shared ggml-cuda sources evaluated to 0.  Kernel bodies gated on
the architecture therefore compiled to nothing, for example the q8_0 -> f16
dequantization kernel in convert.cu, whose NO_DEVICE_CODE fallback expands to
an empty body in host code.

Report the newest architecture like the HIP backend does and exclude the
NVIDIA-only features explicitly, as they are not usable on MUSA.  Define it
for device passes only: CUB uses defined(__CUDA_ARCH__) to detect device
compilation, which is also how nvcc behaves.

Drop the now-redundant defined(__CUDA_ARCH__) checks in the architecture
comparisons: __CUDA_ARCH__ is undefined in host passes for CUDA and MUSA, and
HIP defines it for every pass, so both forms select the same branch.
2026-10-06 10:38:11 +03:00
Captain-Tripps d03bfd7f3c SYCL: reduce tensor allreduce sync with pinned host buffers (llama/29604) 2026-10-06 10:38:11 +03:00
Aaron Teo 66e8b957e7 ggml-zdnn: impl buffer reset, fix memory leaks (llama/29637)
cont: fix code style

Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
2026-10-06 10:38:11 +03:00
cqderekandMax Krasnyansky b8051db82a Hexagon f16 activation ops (llama/29209)
* hexagon: add F16 support for activation ops (SILU/GELU/GELU_QUICK/GEGLU/SWIGLU)

Widens ggml_hexagon_supported_activations() to accept F16 (src0/dst/src1
must agree on type), and adds F16 per-thread worker functions in
act-ops.c mirroring the existing F32 workers, backed by new HVX f16
kernels (hvx_sigmoid_f16_aa, hvx_tanh_f16_aa, hvx_mul_mul_f16_aa,
hvx_min_scalar_f16 family).

SILU, GELU, GELU_QUICK, GEGLU, and SWIGLU are verified correct on-device
(QRD8850) via test-backend-ops CPU-diffed correctness tests. SWIGLU_OAI's
F16 path is code-complete and builds clean on host + all 4 DSP arch
variants (v73/v75/v79/v81), but has no F16 test-case coverage in
test-backend-ops and is therefore unverified on-device in this change.

* hex-ops: align macros

* hex-ops: minor formatting

---------

Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
2026-10-06 10:38:11 +03:00
Xiang Chen 6e48a365f3 gguf : reject tensor size that wraps after padding (llama/26979)
GGML_PAD(nbytes, alignment) wraps to 0 when nbytes is within
(alignment - 1) of SIZE_MAX, which silently bypassed the size
overflow guard in gguf_init_from_reader. Reject the tensor before
padding when nbytes + (alignment - 1) would overflow.

Adds a test-gguf handcrafted case (F32, ne = [4, 2^30-1, 2^30+1, 1])
whose ggml_nbytes = 2^64 - 16 lands in the wrap window. Fails on
master, passes with the guard.
2026-10-06 10:38:11 +03:00
Adrien Gallouët 14423995ba ggml : check row bounds in get_rows_back (llama/29575)
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
2026-10-06 10:38:10 +03:00
Trivikram Reddy 8d2a2cb471 hexagon: optimize concat op (llama/29673)
* hex-concat: reduce pkts in gather/transpose hot loop

gather directly into dst buffer, use special instruction for gather sync

* hex-concat: use fastdiv

replace calls to sw divide with fastpath

* hex-concat: optimize DMA-HVX pipeline and add transpose helpers
2026-10-06 10:38:10 +03:00
Pascal 45093fc8f0 CUDA: bitonic argsort handles rows wider than one block (llama/28957)
Without CUB (HIP, MUSA) argsort ran the bitonic kernel with one thread
per padded column, so any row above 1024 entries launched an invalid
block configuration. Each thread now owns several columns, every stage
of the network runs all owned columns before the barrier, and the block
is capped at 1024 threads. Shared memory becomes the only bound, which
supports_op checks against the device instead of a fixed 1024.

Rows up to 1024 run the same work as before. Bit-exact with the CUB
path on rows of 2048.
2026-10-06 10:38:10 +03:00
thelittlefiremanandCarl Philipp Klemm 482baacdad ggml-cuda: HIP: optimize packed byte subtraction (__vsubss4 -> __vsub4) (llama/29478)
* ggml-cuda: HIP: optimize non-saturating packed byte subtraction (`__vsubss4`)

* CI: ignore 1 spilled vgpr in fattn_vec

---------

Co-authored-by: Carl Philipp Klemm <carl@uvos.xyz>
2026-10-06 10:38:10 +03:00