whisper.cpp

History

Adrien Gallouët 316d921c1a ggml : fix AMX and add batched support (llama/19925) llama-perplexity -hf ggml-org/Qwen3-0.6B-GGUF:Q4_0 -f wikitext-2-raw/wiki.test.raw -c 2048 -b 2048 --chunks 2 before this commit: ``` perplexity: calculating perplexity over 2 chunks, n_ctx=2048, batch_size=2048, n_seq=1 perplexity: 2.31 seconds per pass - ETA 0.07 minutes [1]17.3868,[2]22.2199, Final estimate: PPL = 22.2199 +/- 1.59692 llama_perf_context_print: load time = 878.56 ms llama_perf_context_print: prompt eval time = 2037.82 ms / 4096 tokens ( 0.50 ms per token, 2009.99 tokens per second) llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second) llama_perf_context_print: total time = 6403.17 ms / 4097 tokens llama_perf_context_print: graphs reused = 0 llama_memory_breakdown_print: \| memory breakdown [MiB] \| total free self model context compute unaccounted \| llama_memory_breakdown_print: \| - Host \| 845 = 318 + 224 + 302 \| llama_memory_breakdown_print: \| - CPU_REPACK \| 288 = 288 + 0 + 0 \| llama_memory_breakdown_print: \| - AMX \| 31 = 31 + 0 + 0 \| ``` after this commit: ``` perplexity: calculating perplexity over 2 chunks, n_ctx=2048, batch_size=2048, n_seq=1 perplexity: 1.98 seconds per pass - ETA 0.05 minutes [1]17.2005,[2]21.8220, Final estimate: PPL = 21.8220 +/- 1.56485 llama_perf_context_print: load time = 719.23 ms llama_perf_context_print: prompt eval time = 1676.23 ms / 4096 tokens ( 0.41 ms per token, 2443.58 tokens per second) llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second) llama_perf_context_print: total time = 4258.74 ms / 4097 tokens llama_perf_context_print: graphs reused = 0 llama_memory_breakdown_print: \| memory breakdown [MiB] \| total free self model context compute unaccounted \| llama_memory_breakdown_print: \| - Host \| 845 = 318 + 224 + 302 \| llama_memory_breakdown_print: \| - AMX \| 319 = 319 + 0 + 0 \| ``` (no more CPU_REPACK) after this commit, disabling amx: ``` perplexity: calculating perplexity over 2 chunks, n_ctx=2048, batch_size=2048, n_seq=1 perplexity: 2.34 seconds per pass - ETA 0.07 minutes [1]17.2005,[2]21.8220, Final estimate: PPL = 21.8220 +/- 1.56485 llama_perf_context_print: load time = 841.91 ms llama_perf_context_print: prompt eval time = 2057.28 ms / 4096 tokens ( 0.50 ms per token, 1990.98 tokens per second) llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second) llama_perf_context_print: total time = 6454.51 ms / 4097 tokens llama_perf_context_print: graphs reused = 0 llama_memory_breakdown_print: \| memory breakdown [MiB] \| total free self model context compute unaccounted \| llama_memory_breakdown_print: \| - Host \| 845 = 318 + 224 + 302 \| llama_memory_breakdown_print: \| - CPU_REPACK \| 319 = 319 + 0 + 0 \| ``` => same perplexity. Signed-off-by: Adrien Gallouët <angt@huggingface.co>		2026-02-27 20:57:58 +02:00
..
ggml-blas	ggml : add ggml_build_forward_select (llama/18550)	2026-01-30 15:56:40 +02:00
ggml-cann	CANN: Remove unnecessary wrapper for `gml_backend_buft_is_cann` (llama/18968)	2026-02-15 21:44:37 +02:00
ggml-cpu	ggml : fix AMX and add batched support (llama/19925)	2026-02-27 20:57:58 +02:00
ggml-cuda	Improve CUDA graph capture (llama/19754)	2026-02-27 20:57:58 +02:00
ggml-hexagon	hexagon refactor all Ops to use local context struct (llama/19819)	2026-02-27 20:57:58 +02:00
ggml-hip	HIP: add mmf for CDNA (llama/18896)	2026-01-30 15:56:40 +02:00
ggml-metal	models : optimize qwen3next graph (llama/19375)	2026-02-15 21:44:37 +02:00
ggml-musa	…
ggml-opencl	opencl: refactor expm1 and softplus (llama/19404)	2026-02-27 20:57:58 +02:00
ggml-rpc	rpc : use unordered_map::reserve and emplace (llama/18513)	2026-01-14 09:11:59 +02:00
ggml-sycl	support permuted, remove check s0/s10 (llama/19889)	2026-02-27 20:57:58 +02:00
ggml-virtgpu	ggml-virtgpu: improve the reliability of the code (llama/19846)	2026-02-27 20:57:58 +02:00
ggml-vulkan	vulkan: fix fp16 Flash Attention on Windows AMD RDNA2 and below (llama/19921)	2026-02-27 20:57:58 +02:00
ggml-webgpu	ggml-webgpu: Add unary op (SQR, SQRT, SIN, COS) support. (llama/19700)	2026-02-27 20:57:58 +02:00
ggml-zdnn	ggml-zdnn : mark zDNN buffers as non-host (llama/18967)	2026-01-30 15:56:40 +02:00
ggml-zendnn	ggml-zendnn : resolve ZenDNN backend cross-module symbol dependency (llama/19159)	2026-01-30 15:56:40 +02:00
CMakeLists.txt	hexagon: enable offloading to Hexagon on Windows on Snapdragon (llama/19150)	2026-01-30 15:56:40 +02:00
ggml-alloc.c	ggml : make `ggml_is_view` as API (llama/19539)	2026-02-27 20:57:58 +02:00
ggml-backend-dl.cpp	hexagon: enable offloading to Hexagon on Windows on Snapdragon (llama/19150)	2026-01-30 15:56:40 +02:00
ggml-backend-dl.h	hexagon: enable offloading to Hexagon on Windows on Snapdragon (llama/19150)	2026-01-30 15:56:40 +02:00
ggml-backend-impl.h	llama: use host memory if device reports 0 memory (llama/18587)	2026-01-14 09:11:59 +02:00
ggml-backend-reg.cpp	ggml : use noexcept overload for is_regular_file in backend registration (llama/19452)	2026-02-15 21:44:37 +02:00
ggml-backend.cpp	ggml-backend: fix async set/get fallback sync (llama/19179)	2026-02-08 09:29:10 +02:00
ggml-common.h	…
ggml-impl.h	ggml : make `ggml_is_view` as API (llama/19539)	2026-02-27 20:57:58 +02:00
ggml-opt.cpp	…
ggml-quants.c	…
ggml-quants.h	…
ggml-threading.cpp	…
ggml-threading.h	…
ggml.c	ggml/gguf : prevent integer overflows (llama/19856)	2026-02-27 20:57:58 +02:00
ggml.cpp	…
gguf.cpp	…