whisper.cpp/ggml
Masashi Yoshimura 2955e84ec3 ggml-webgpu: improve MTP inference by using mat-vec path for small batches (llama/24811)
* ggml-webgpu: improve small batches decoding

* Add barrier to the NUM_COLS loop in mul-mat-vec
2026-06-26 16:03:57 +03:00
..
cmake ggml : Parallelize quant LUT init (llama/23595) 2026-05-25 12:26:07 +03:00
include Remove padding and multiple D2D copies for MTP (llama/24086) 2026-06-15 10:33:53 +03:00
src ggml-webgpu: improve MTP inference by using mat-vec path for small batches (llama/24811) 2026-06-26 16:03:57 +03:00
.gitignore whisper : reorganize source code + improve CMake (#2256) 2024-06-26 19:34:09 +03:00
CMakeLists.txt ggml : bump version to 0.15.2 (ggml/1548) 2026-06-19 12:53:43 +03:00