Commit Graph
5022 Commits
Author SHA1 Message Date
David Friehs 9d8e6b91b1 cuda: unblock mmq for MoE on sm_60 (llama/26264) 2026-09-04 13:39:29 +03:00
8c0adb05fe ggml-metal: add chunked SSD MMA for Mamba-2 prefill optimization (llama/26647)
* metal: WIP chunked SSD SSM_SCAN kernels for multi-token prefill

* metal: drop scalar SSD path; MMA + sequential tail

* drop WIP ssm scan test noise

* remove state_from_dst and rename CS and NSG constants

* remove unrelated  added whitespace padding

* added clarity to mma_tokens calculation

* added clarity to use_mma bool checks

* added comments to metal ssd op constants for clarity

* reserve K tokens for sequential kernel rollback snapshots

* reset concurrency between mma and seq tail

* remove print args no longer used

* fixed comment to no longer point to specific line

* add FC_SSM_SCAN so seq path skips token offlset unless it's mma tail

* added changes to new ssm.metal for rebase after ggml-metal.metal refactor

* specialize ssm_scan tail with a template instead of a function constant

---------

Co-authored-by: dpantaleoni <dominikpantaleoni@gmail.com>
Co-authored-by: forforever73 <690105611@qq.com>
2026-09-04 13:39:29 +03:00
Max Krasnyansky 8df657a2fa ggml-meta: propagate buffer usage and call init on the new tensors (llama/27586) 2026-09-04 13:39:29 +03:00
Jonathan Clohessy 482956e744 kleidiai: Rework KleidiAI Build System/Integration (llama/26077)
* Rework KleidiAI Build System/Integration

Signed-off-by: Jonathan Clohessy <Jonathan.Clohessy@arm.com>

* Add fp16 guard, and fix cmake caching issue

Signed-off-by: Jonathan Clohessy <Jonathan.Clohessy@arm.com>

* Fix formatting, and rebase issue

Signed-off-by: Jonathan Clohessy <Jonathan.Clohessy@arm.com>

---------

Signed-off-by: Jonathan Clohessy <Jonathan.Clohessy@arm.com>
2026-09-04 13:39:29 +03:00
Ryan CandRyan Churaman e820c280a6 rpc: support apple RDMA as an RPC transport (llama/26421)
* rpc: support apple RDMA as an RPC transport

* remove set_tensor micro optimization, rpc socket pinning per CR

* remove transparent reconnect

* trigger apple builds on RPC changes

---------

Co-authored-by: Ryan Churaman <rschu@meta.com>
2026-09-04 13:39:28 +03:00
Yuri KhrustalevandGeorgi Gerganov be12d39a17 metal : null-check buffer alloc to fix OOM crash (llama/25371)
* metal : null-check ggml_metal_buffer_init result to avoid OOM crash

ggml_backend_metal_buffer_type_alloc_buffer used the result of
ggml_metal_buffer_init without checking for NULL. ggml_metal_buffer_init
returns NULL when the underlying Metal allocation fails (e.g. an
out-of-memory condition), and the following ggml_metal_buffer_is_shared(res)
call dereferences it, turning a recoverable allocation failure into a hard
crash (EXC_BAD_ACCESS). This is easy to hit on memory-constrained devices
such as iOS when a model/context exceeds the available Metal budget.

Log the failure using the existing GGML_LOG_ERROR convention and return
NULL so the allocator surfaces a diagnosable error up the stack instead of
crashing.

* cont : fix log

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
2026-09-04 13:39:28 +03:00
KITAITI MakotoandDaniel Bevenius 642b5d3260 ruby : Add #free method, check MemoryView strictly (#4032)
* Bump version to 1.3.9

* Fix JFKReader's format

* Add Whisper::Context#free

* Add test for Whisper::Context#free

* Add Whisper::Context#free to RBS

* Add Whisper::VAD::Context#free

* Add test for Whisper::VAD::Context#free

* Add Whisper::VAD::Context#free to RBS

* Add Parakeet::Context#free

* Add test for Parakeet::Context#free

* Add Parakeet::Context#free to RBS

* Check MemoryView format more strictly

* Fallback MemoryView flags to SIMPLE when MemoryView not available

* Add changelog entries

* Update document on callbacks

* Remove whitespace

Co-authored-by: Daniel Bevenius <daniel.bevenius@gmail.com>

---------

Co-authored-by: Daniel Bevenius <daniel.bevenius@gmail.com>
2026-09-04 14:33:37 +09:00
Ben Younes eacbd8234c whisper : default-initialize whisper_mel to avoid uninitialized read (#3981)
whisper_full()/whisper_full_with_state() only compute the mel spectrogram when
n_samples > 0. For n_samples == 0 on a freshly allocated state the mel is never
touched, but struct whisper_mel had no member initializers, so n_len / n_len_org
/ n_mel were indeterminate heap garbage (the state is allocated with
new whisper_state). seek_end is derived from that garbage and, depending on it,
the call either quietly returns 0 or runs the encoder with garbage dimensions
over an empty (NULL) mel buffer, dereferencing address 0 in the mel copy loop.

Give whisper_mel default member initializers so a never-computed mel reads as
0 frames and the n_samples == 0 case deterministically takes the existing
too-short path.

Fixes #3978
2026-08-31 05:53:57 +02:00
Jhen-Jie Hong c4ac0012a8 parakeet : fix TDT decode by outputting raw logits from the joint graph (#4017) 2026-08-29 05:28:05 +02:00
Georgi Gerganov 978113305b talk-llama : sync llama.cpp 2026-08-25 15:28:56 +03:00
Georgi Gerganov 3680f66fa4 pi : init 2026-08-25 15:28:56 +03:00
Georgi Gerganov 0414519d7f sync : ggml 2026-08-25 15:28:56 +03:00
Georgi Gerganov d470c9d469 ggml : bump version to 0.22.0 (ggml/1607)
* ggml : bump version to 0.22.0

* scripts : update default release desc
2026-08-25 15:28:56 +03:00
Neo Zhang 322a77cf6d sycl : mark tq2_0 as not supported (llama/27660) 2026-08-25 15:28:56 +03:00
fairydreamingandStanisław Szymczyk 17a522a8cc webgpu : fix handling of infinity values during ARGSORT and TOP_K (llama/27538)
Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
2026-08-25 15:28:56 +03:00
YiChen LvandGeorgi Gerganov fa3b87c188 metal : per-device tuned (Q, NE) for flash-attn vec (llama/26570)
* metal : per-device tuned (Q, NE) for flash-attn vec (llama/25750)

* rebase Q-generic FA vec body from 01dc93607 (llama/23114)

* add 53 f16 (Q,NE) flash-attn vec instantiations (vec 80 -> 133)

* add FA vec (Q,NE) tuning table + dispatch wiring + SMEM cap fallback

* add  FA vec (Q,NE) perf sweep

* fill tuning result

* fold family table into a per-family representative SKU

* refactor tuning result format

* extend FA vec tuning to quantized KV caches

* sync fa vec tuner bucketing with runtime, use pointwise tuning regret

* update tuned table

* format and cleanup

* prefix fa_vec tuning procs with ggml_backend_metal_tuning_, drop unused fa_vec_override_active

* add device id -> token lookup for the offline tuning tool

* add ggml-metal-tuning skeleton

* add op-agnostic perf cell + median timing for the tuner

* add FA-vec graph build + tensor init to the tuner

* tools : add FA-vec (Q,NE) sweep, compression and table emit

* cool down and re-measure the dirty window on thermal drift

* test-backend-ops : replace the FA vec tune mode with a bounded (Q,NE) slice

* tools : document the Metal tuner, point the table comment at it

* abort on unknown KV type, single-source fa_vec_legal_ne

* cleanup

* honor -o in the FA vec (Q,NE) slice

* retune FA-vec (Q, NE) under a pointwise no-harm gate

* cont : add fa-vec tunings for M1 Pro, M2 Ultra, M5 Max

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
2026-08-25 15:28:56 +03:00
Georgi Gerganov fe52277c1f sync : ggml 2026-08-25 15:28:56 +03:00
YiChen Lv b15d31d0fd metal: per-op source split + parallel compile (llama/26561) 2026-08-25 15:28:56 +03:00
Georgi Gerganov 4257445149 scripts : update ggml-am 2026-08-25 15:28:56 +03:00
Georgi Gerganov 1d8e05291a sync : ggml 2026-08-25 15:28:56 +03:00
Georgi Gerganov aa25d33f58 ggml : shorten virtual device naming in CUDA and Metal (llama/27608)
* ggml : shorten virtual device naming in CUDA and Metal

Assisted-by: llama.cpp:DeepSeek-V4-Flash-0731

* ggml-metal : build device description at init

Assisted-by: llama.cpp:DeepSeek-V4-Flash-0731

* cont : naming
2026-08-25 15:28:56 +03:00
fairydreamingandStanisław Szymczyk 103305e656 webgpu : reorder includes since V that appears in common_decls.tmpl may be defined as K in flash_attn_decls.tmpl if KV_OVERLAP (llama/27545)
Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
2026-08-25 15:28:56 +03:00
Georgi Gerganov 32d0f49ddc ggml : fix ggml_clamp (llama/27644)
* ggml : fix ggml_clamp

* cont : update ggml-alloc
2026-08-25 15:28:56 +03:00
Aman Gupta 20209c25f7 Deepseek 4: -sm tensor (llama/26490)
* DSV4: sm tensor

* set coarser granularity for head splits

* fix dspark

* add model saving for dsv4 + allow dflash to return on specific device

* add comment about dsv4 seq_rm

* simplify

* add shared expert delayed allreduce

* remove special test for dsv4
2026-08-25 15:28:56 +03:00
Gaurav Garg 3b89b37c30 Fix meta tensor split state propagation (llama/27574)
* ggml : fix meta tensor split state propagation

* Add test-llama-archs to CI
2026-08-25 15:28:56 +03:00
Aman Karki c8a40099fa cuda : add POOL_1D support (llama/27573)
* cuda : add POOL_1D support

* fix: add missing trailing newline for editorconfig compliance
2026-08-25 15:28:56 +03:00
Safi UllahandJeff Bolz 13a7856daa vulkan : added the PAD_REFLECT_1D operation (llama/26586)
* vulkan : added PAD_REFLECT_1D operation

Implemented the GGML_OP_PAD_REFLECT_1D operation for the Vulkan backend

Changes:
- pad_reflect_1d.comp: implemented the GLSL compute shader with reflection logic
- vulkan-shaders-gen.cpp: register the shader for SPIR-V compilation
- ggml-vulkan.cpp: pushed constants struct, pipeline creation,
  supports_op, dispatch function, compute switch and debug validation

Tested the PAD_REFLECT_1D on Intel Iris Xe (Vulkan 1.4, Mesa 25.2.8):

Correctness:
  PAD_REFLECT_1D(type=f32,ne_a=[512,34,2,1],pad_0=10,pad_1=9) = Pass
  PAD_REFLECT_1D(type=f32,ne_a=[3000,384,4,1],pad_0=10,pad_1=9) = Pass
  2/2 tests passed
 - All test are passed

Performance:
  ne_a=[512,34,2,1] -> 5.38 us/run, 24.55 GB/s
  ne_a=[3000,80,1,1] -> 30.09 us/run, 59.62 GB/s
  ne_a=[3000,384,4,1] -> 158.31 us/run, 54.39 GB/s

* Update ggml/src/ggml-vulkan/vulkan-shaders/pad_reflect_1d.comp

Co-authored-by: Jeff Bolz <jbolz@nvidia.com>

---------

Co-authored-by: Jeff Bolz <jbolz@nvidia.com>
2026-08-25 15:28:56 +03:00
Kartik SirohiandSigbjørn Skjæret 1efb31e6fd ggml: optimize concat op by replacing per-element memcpy with row-level memcpy (llama/24575)
* ggml: optimize concat op by replacing per-element memcpy with row-level memcpy

* ggml: fix concat offsets for row-level copies

* ggml: add concat row contiguity asserts

* ggml: move concat block size asserts

* ggml: remove redundant concat asserts

* Update ggml/src/ggml-cpu/ops.cpp

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>

---------

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
2026-08-25 15:28:56 +03:00
Sigbjørn Skjæret 21a67dd8da sycl : add Q2_K reordered MMVQ and ESIMD kernels (again) (llama/27490)
* Revert "Revert "sycl : add Q2_K reordered MMVQ and ESIMD kernels (#26336)" (#…"

This reverts commit 7a0e42fd01fb0acda644e4f04b1f1acbbb9e23ba.

* add gate params
2026-08-25 15:28:56 +03:00
Hongqiang Wang 5f7bd9ddb9 opencl: fold the gpt-oss MoE per-expert bias adds into the epilogue (op/kernel fusion) (llama/26431)
* opencl: fold the gpt-oss MoE bias adds into swiglu_oai

Default on, opt out with GGML_OPENCL_FUSE_MOE_BIAS_GLU=0.

* opencl: fold the MoE down-projection bias into the combine

Default on, opt out with GGML_OPENCL_FUSE_MOE_BIAS_COMBINE=0.
2026-08-25 15:28:56 +03:00
Ben YounesandBen Younes a722846cb6 whisper : guard null source in buffer loader read callback (#3982)
* whisper : guard null source in buffer loader read callback

whisper_init_from_buffer_with_params_no_state installs a read callback that
copies from buf->buffer + current_offset. When the buffer is exhausted (or the
supplied buffer is empty), size_to_copy is 0 and the source pointer can be null;
passing a null pointer to memcpy is undefined behavior even for a zero-length
copy (UBSan: 'null pointer passed as argument 2' at the memcpy). Loading a
crafted/short model through the buffer loader could hit this.

Skip the memcpy when there is nothing to copy. Loading from a null/empty or
truncated buffer now fails gracefully (returns NULL) with no UB.

This addresses bug 1 of #3879. Bug 2 (integer overflow when sizing the mel
filter buffer) is covered by the open PR #3780.

* fixup! whisper : guard null source in buffer loader read callback

---------

Co-authored-by: Ben Younes <2910651+ousamabenyounes@users.noreply.github.com>
2026-08-25 12:40:19 +02:00
Daniel Bevenius c122757fdd docs : center badges in README.md [no ci] (#4012) 2026-08-24 12:36:17 +02:00
berney 25694098b8 devops : add main-rocm Dockerfile (#3975) 2026-08-24 11:40:24 +02:00
52dec9d889 vitisai : add VitisAI Plugin for AMD Ryzen AI NPU encoder offload (#3608)
* Add VitisAI Plugin

* Added VitisAI encoder module placeholder files

* VitisAI build integration

* VitisAI encoder offload functional

* Clean up vitisai integration

* Add c++17 requirement for Windows

* Enabled preemption for windows runs

* Add model cache override option

* Remove vitisai premature log message

* Add rai support through file mapping

* Fixed flatbuffer loading

* Fixed Windows file mapping issue

* Update FlexmlRT resolution

* Use Flexmlrt wheel pkg to build VitisAI plugin

* Clean up

* Remove prints

* Change flexmlrt target from Shared to Interface

* Add c++17 requirement for Windows

* Enabled preemption for windows runs

* Add rai support through file mapping

* Fixed flatbuffer loading

* Fixed Windows file mapping issue

* Update FlexmlRT resolution

* Use Flexmlrt wheel pkg to build VitisAI plugin

* Clean up

* Remove prints

* Change flexmlrt target from Shared to Interface

* Cleanup FlexmlRT integration

* format fix

* Adding AMD Licenses

* Update CMakeLists.txt

Co-authored-by: Kumawat, Sachin <sachin.kumawat@amd.com>

* Update src/CMakeLists.txt

Co-authored-by: Kumawat, Sachin <sachin.kumawat@amd.com>

* Update whisper.cpp

* Added VitisAI encoder readme section

* Remove license headers from common files to whisper.cpp

---------

Co-authored-by: Sachin Kumawat <sachink@amd.com>
Co-authored-by: Jeff Lin <jeffylin@xilinx.com>
Co-authored-by: Lin <jefflin@amd.com>
Co-authored-by: Lin, Jeff (DCG-ENG) <jeff.lin@amd.com>
Co-authored-by: Iswarya Alex <iswaryaalex96@gmail.com>
Co-authored-by: Alex, Iswarya <Iswarya.Alex@amd.com>

* Update README.md

- RAI EULA Links
- Updated for RAI Whisper instructions

* Cleanup and add runtime print debug guard

* turn off profiling

* Add VitisAI Plugin

* Added VitisAI encoder module placeholder files

* VitisAI build integration

* VitisAI encoder offload functional

* Clean up vitisai integration

* Add c++17 requirement for Windows

* Enabled preemption for windows runs

* Add model cache override option

* Remove vitisai premature log message

* Add rai support through file mapping

* Fixed flatbuffer loading

* Fixed Windows file mapping issue

* Update FlexmlRT resolution

* Use Flexmlrt wheel pkg to build VitisAI plugin

* Clean up

* Remove prints

* Change flexmlrt target from Shared to Interface

* Add c++17 requirement for Windows

* Enabled preemption for windows runs

* Add rai support through file mapping

* Fixed flatbuffer loading

* Fixed Windows file mapping issue

* Update FlexmlRT resolution

* Use Flexmlrt wheel pkg to build VitisAI plugin

* Clean up

* Remove prints

* Change flexmlrt target from Shared to Interface

* Cleanup FlexmlRT integration

* format fix

* Adding AMD Licenses

* Update CMakeLists.txt

Co-authored-by: Kumawat, Sachin <sachin.kumawat@amd.com>

* Update src/CMakeLists.txt

Co-authored-by: Kumawat, Sachin <sachin.kumawat@amd.com>

* Update whisper.cpp

* Added VitisAI encoder readme section

* Remove license headers from common files to whisper.cpp

---------

Co-authored-by: Sachin Kumawat <sachink@amd.com>
Co-authored-by: Jeff Lin <jeffylin@xilinx.com>
Co-authored-by: Lin <jefflin@amd.com>
Co-authored-by: Lin, Jeff (DCG-ENG) <jeff.lin@amd.com>
Co-authored-by: Iswarya Alex <iswaryaalex96@gmail.com>
Co-authored-by: Alex, Iswarya <Iswarya.Alex@amd.com>

* Cleanup and add runtime print debug guard

* Update README.md

- RAI EULA Links
- Updated for RAI Whisper instructions

* turn off profiling

* Let flexmlrt detect device type

* Add VitisAI model download scripts

* Add encoder + cross projection layer offload

* Add self hosted runner for amd npu

* Update runner

* Update workflow for linux

* Update workflow for linux

* Update flexmlrt packages for linux

* Update flexmlrt packages for linux

* Updated README

* readme: clarify xrt

* readme: clarify xrt

* ci: update test config

* Added supported plarform details with python 3.12 requirement for Linux

* Use refactored helpers

* Deprecate cross_proj .rai naming and cleanup

* Remove stale function code

* Fix: formatting

---------

Co-authored-by: Jeff Lin <jeffylin@xilinx.com>
Co-authored-by: Lin <jefflin@amd.com>
Co-authored-by: Lin, Jeff (DCG-ENG) <jeff.lin@amd.com>
Co-authored-by: Iswarya Alex <iswaryaalex96@gmail.com>
Co-authored-by: Alex, Iswarya <Iswarya.Alex@amd.com>
Co-authored-by: Iswarya Alex <47045679+iswaryaalex@users.noreply.github.com>
2026-08-24 08:02:30 +02:00
Mostafa 233fe1fc9b whisper : bypass cross-attention scaling for OpenVINO backend (#3997)
* whisper : bypass cross-attention scaling for OpenVINO backend

* whisper : use correct gating flag
2026-08-22 07:14:11 +02:00
Daniel Bevenius 51de5e8bb0 openvino : update model conversion and README.md (#4003)
This commit updates the OpenVINO model conversion script and its
dependencies as they were currently not working.

The first issue was an import that needed updating for fix the following
error:
```console
  (openvino_conv_env) $ python convert-whisper-to-openvino.py --model base.en
  Traceback (most recent call last):
    File "/home/danbev/work/ai/whisper-work/models/convert-whisper-to-openvino.py", line 6, in <module>
      from openvino.runtime import serialize
  ModuleNotFoundError: No module named 'openvino.runtime'
```
And after that there was a missing dependency:
```console
ModuleNotFoundError: No module named 'onnxscript'
```

With the changes in this commit I was able to successfully convert the model
and run the inference using OpenVINO.
2026-08-22 07:13:26 +02:00
Daniel Bevenius 3391d6b6f9 scripts : add release.sh script (#4010)
* scripts : add release.sh script [no ci]

This commit adds a release.sh script similar to what ggml and llama.cpp
have to prepare a release.

* scripts : use sed_inplace [no ci]
2026-08-22 07:12:42 +02:00
Daniel Bevenius a4610c78bf docs : add release badge and remove stable/roadmap [no ci] (#4009)
This commit updates the README.md to include a release badge.

The motivation for this is that we have been manually updating the
Stable link in this document for releases. Removing this means that we
don't have to touch this file and only update CMakeLists.txt.

I also removed the link to the Roadmap was it felt out of place and it
has not been kept up to date so I hope that is alright.

I'll follow up with a PR to introduce a release script that will help
assist the creation of releases (updateing the version and creating the
release PR etc).
2026-08-22 07:12:05 +02:00
Daniel Bevenius ab578879ec make : add --parallel to cmake build command (#4007)
This commit adds the --parallel flag to the cmake build command to
speed up compilation. This will hopefully help a little with CI runs
even though the are often limited to 2 cores.
2026-08-22 07:11:20 +02:00
Georgi Gerganov 45f1593fd3 sync : ggml 2026-08-21 19:23:26 +03:00
Georgi Gerganov ce77728c06 Revert "sycl : add Q2_K reordered MMVQ and ESIMD kernels (llama/26336)" (llama/27486)
This reverts commit ff14356e0caf6988f61f1f15f9dfe7d5ab398271.
2026-08-21 19:23:26 +03:00
Georgi Gerganov 0d9ba28cde ggml : bump version to 0.21.0 (ggml/1597) 2026-08-21 19:23:26 +03:00
Charles Xu d6c416e222 kleidiai : add SME2 F32 GEMV kernel support (llama/26891) 2026-08-21 19:23:26 +03:00
Todd Malsbary d60ef65077 sycl : add Q2_K reordered MMVQ and ESIMD kernels (llama/26336)
* Add DMMV Q4_K and Q6_K ESIMD kernels

Configure cmake build with -DGGML_SYCL_ESIMD=ON to enable.

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

* Refactor ESIMD kernels to share common code

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

* Move control of ESIMD from compile to runtime

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

* Use ESIMD by default when available

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

* Fix possible error when using ESIMD by default

While not an issue in the current version, this will become an
issue when additional QK ESIMD kernels are added (such as Q2_K).

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

* Add explicit unroll to ESIMD kernels

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

* Tidy up ESIMD kernels a bit

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

* Add a reordered Q2_K MMVQ kernel

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

* Add DMMV Q2_K ESIMD kernel

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

---------

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>
2026-08-21 19:23:26 +03:00
Todd Malsbary b1cb805965 sycl : Add Q5_K ESIMD kernel (llama/26376)
* Add DMMV Q4_K and Q6_K ESIMD kernels

Configure cmake build with -DGGML_SYCL_ESIMD=ON to enable.

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

* Refactor ESIMD kernels to share common code

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

* Move control of ESIMD from compile to runtime

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

* Use ESIMD by default when available

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

* Fix possible error when using ESIMD by default

While not an issue in the current version, this will become an
issue when additional QK ESIMD kernels are added (such as Q2_K).

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

* Add explicit unroll to ESIMD kernels

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

* Tidy up ESIMD kernels a bit

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

* Add DMMV Q5_K ESIMD kernel

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

* Remove redundant copyright notice

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

---------

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>
2026-08-21 19:23:26 +03:00
Hongqiang WangandLi He 2cb52ddbc0 opencl: keep the vocab-scale K-quant lm_head on the CPU for Adreno A7X (compiler issue workaround) (llama/26440)
* opencl: keep the vocab-scale K-quant lm_head on the CPU on the Adreno A7X

* opencl: revise comments

---------

Co-authored-by: Li He <lih@qti.qualcomm.com>
2026-08-21 19:23:26 +03:00
HumerousGorgonandNeo Zhang af74f97965 sycl: Update gate logic for Alchemist GPUs regarding OneDNN features. (llama/26635)
* feat: updated gating logic of fattn-onednn.cpp

* verified device types

* Update ggml/src/ggml-sycl/fattn-onednn.cpp

Accepted recommendations to add bmg_g31 arch.

Co-authored-by: Neo Zhang <zhang.jianyu@outlook.com>

* Improved SPDA gate, added documentation.

* Added arch var to reworked gate, fixing build errors.

* Fix trailing whitespaces.

---------

Co-authored-by: Neo Zhang <zhang.jianyu@outlook.com>
2026-08-21 19:23:26 +03:00
Ian Faust f19250e5ac sycl: fix multiple warnings in compiling sycl backend (llama/26713)
* Update norm.cpp

* Update helper.hpp

* Update im2col.cpp

* Update fattn-mkl.cpp

* Update element_wise.cpp

* Update fattn-mkl.cpp

* Update set_rows.cpp

* Update element_wise.cpp

* Update ggml-sycl.cpp

* Update ggml-sycl.cpp

* Update ggml-sycl.cpp

* Update ggml-sycl.cpp

* Update ggml-sycl.cpp

* Update norm.cpp

* Update CMakeLists.txt

* Update CMakeLists.txt

* Update CMakeLists.txt

* Update ggml-sycl.cpp
2026-08-21 19:23:26 +03:00
Neo Zhang 12137c316d sycl : fix load model with mlock issue (llama/27250) 2026-08-21 19:23:26 +03:00
Xuan-Son Nguyen 5656e44ea2 ggml: support ggml_rope_set_offset on opencl, sycl, wgpu, hexagon (llama/27345)
* ggml: support ggml_rope_set_offset on opencl, sycl, wgpu, hexagon

* rm inplace optimization
2026-08-21 19:23:26 +03:00