Vulkan: route large matmuls to medium tile on Adreno (llama/24877)

* [Vulkan] Fixes llama-cli breaking over longer promts sizes

The llama-cli was breaking for longer promts sizes for q4_0 quantized networks. Causing due to insufficient shared memory.

* Removed the un-used Adreno device

* Updated matmul for small pipeline.
This commit is contained in:
Raman Shinde 2026-07-11 13:58:29 +05:30 committed by Georgi Gerganov
parent ea8085373a
commit 4bc66b579f
No known key found for this signature in database
GPG Key ID: 449E073F9DC10735
1 changed files with 8 additions and 0 deletions

View File

@ -6501,6 +6501,14 @@ static vk_device ggml_vk_get_device(size_t idx) {
device->mul_mat_id_m[i] = true;
device->mul_mat_id_s[i] = false;
break;
case VK_VENDOR_ID_QUALCOMM:
device->mul_mat_l[i] = false;
device->mul_mat_m[i] = true;
device->mul_mat_s[i] = true;
device->mul_mat_id_l[i] = false;
device->mul_mat_id_m[i] = true;
device->mul_mat_id_s[i] = true;
break;
#endif
default:
device->mul_mat_l[i] = true;