whisper.cpp

History

Chenguang Li dfba84cb47 CANN: support flash attention for head dim not multiple of 16, fix ALiBi slope offset (llama/20031) - Allow FLASH_ATTN_EXT when head dimension D is not a multiple of 16 by padding Q/K/V to D_padded = GGML_PAD(D, 16), running FusedInferAttentionScoreV2, then slicing the output back to D (ggml-cann.cpp + aclnn_ops.cpp). - Fix aclnn_get_slope second-part offset: use ggml_type_size(dtype) instead of sizeof(float) so ALiBi slopes are correct when dtype is F16 (e.g. GQA with 48 heads); fixes buffer overflow and large numerical errors in those cases.		2026-03-29 15:04:36 +03:00
..
cmake	cmake : remove unused file (ggml/1419)	2026-02-08 09:29:10 +02:00
include	ggml : restore ggml_type_sizef() to aboid major version bump (ggml/1441)	2026-03-18 15:18:24 +02:00
src	CANN: support flash attention for head dim not multiple of 16, fix ALiBi slope offset (llama/20031)	2026-03-29 15:04:36 +03:00
.gitignore	…
CMakeLists.txt	ggml : bump version to 0.9.8 (ggml/1442)	2026-03-18 15:18:24 +02:00