Files
whisper.cpp/ggml
Radoslav Gerganov 321f628239 rpc : reuse compute graph buffers (llama/21299)
Reuse the buffer for the ggml context which is used for creating the
compute graph on the server side. This partially addresses a memory leak
created by the CUDA backend due to using buffer addresses as cache
keys.

ref: #21265
ref: #20315
2026-04-30 11:29:00 +03:00
..