whisper.cpp/ggml/include
Johannes Gäßler f8a831779e CUDA: use mma PTX instructions for FlashAttention (llama/11583)
* CUDA: use mma PTX instructions for FlashAttention

* __shfl_sync workaround for movmatrix

* add __shfl_sync to HIP

Co-authored-by: Diego Devesa <slarengh@gmail.com>
2025-02-03 22:00:57 +02:00
..
ggml-alloc.h
ggml-backend.h rpc : early register backend devices (llama/11262) 2025-02-03 22:00:57 +02:00
ggml-blas.h
ggml-cann.h
ggml-cpp.h
ggml-cpu.h
ggml-cuda.h
ggml-kompute.h
ggml-metal.h
ggml-opencl.h
ggml-opt.h
ggml-rpc.h
ggml-sycl.h
ggml-vulkan.h
ggml.h CUDA: use mma PTX instructions for FlashAttention (llama/11583) 2025-02-03 22:00:57 +02:00
gguf.h GGUF: C++ refactor, backend support, misc fixes (skip) (llama/11030) 2025-01-14 10:38:01 +02:00