whisper.cpp/ggml
Hongqiang Wang c4a15ec197
opencl: fix two issues on flash attention for Adreno a7x (llama/25697)
* opencl: route `sub_group_shuffle_xor` to qcom ext when KHR ext is unavailable

KHR `sub_group_shuffle_xor` is not defined by compiler when
`cl_qcom_subgroup_shuffle` is present, causing certain FA
kernels fail to build. Define the KHR shuffle_xor using
the qcom extension.

* opencl: skip FA kernels with mixed and quant types for A7x to avoid compiler crash
2026-07-30 16:31:41 +03:00
..
cmake ggml : Parallelize quant LUT init (llama/23595) 2026-05-25 12:26:07 +03:00
include ggml : add a set of functions for checking contiguity of inner tensor dimensions (llama/25650) 2026-07-30 16:31:39 +03:00
src opencl: fix two issues on flash attention for Adreno a7x (llama/25697) 2026-07-30 16:31:41 +03:00
.gitignore whisper : reorganize source code + improve CMake (#2256) 2024-06-26 19:34:09 +03:00
CMakeLists.txt ggml-et: Initial ET backend (llama/24179) 2026-07-30 16:31:36 +03:00