mirror of
https://github.com/ggml-org/whisper.cpp.git
synced 2026-10-10 08:15:40 +02:00
Reuse the buffer for the ggml context which is used for creating the compute graph on the server side. This partially addresses a memory leak created by the CUDA backend due to using buffer addresses as cache keys. ref: #21265 ref: #20315