mirror of
https://github.com/ggml-org/whisper.cpp.git
synced 2026-10-01 03:59:20 +02:00
Allow HMX flash-attention to run with head_dim not a multiple of 64 (e.g. SigLIP head_dim=72), by operating on DK/DV rounded up to 64 with zero-filled tail lanes.