whisper.cpp/ggml
Oliver Simons 3fee8a1e05 CUDA: Fix data-races when reusing SMEM in block_reduce (llama/26385)
* CUDA: Fix data-races when reusing block_reduce

block_reduce currently doesn't resync after reading from SMEM, causing
potential data-races when reusing SMEM for multiple reductions.

One may consider simply always adding this in block_reduce, but this
comes at a potential perf cost

* double-buffering for single-row softmax

* double-buffering for norm as well

* Add comment

* Add explanatory comment to block_reduce

* Specify need for + do memory barrier only in multi-warp scenario

* Implement review-suggestion from @gaugarg-nv
2026-08-04 13:37:47 +03:00
..
cmake ggml : Parallelize quant LUT init (llama/23595) 2026-05-25 12:26:07 +03:00
include sync : ggml (#3962) 2026-07-31 09:11:28 +02:00
src CUDA: Fix data-races when reusing SMEM in block_reduce (llama/26385) 2026-08-04 13:37:47 +03:00
.gitignore whisper : reorganize source code + improve CMake (#2256) 2024-06-26 19:34:09 +03:00
CMakeLists.txt sync : ggml (#3962) 2026-07-31 09:11:28 +02:00