whisper.cpp

Commit Graph

Author	SHA1	Message	Date
Xuan-Son Nguyen	79fb43e252	ggml : add mrope kernel for metal (llama/13457)	2025-05-13 13:10:08 +03:00
Georgi Gerganov	926e06dbfd	metal : optimize MoE for large batches (llama/13388)	2025-05-13 13:09:20 +03:00
lhez	43a59eccf6	opencl: remove unnecessary assert for `add` (llama/13257)	2025-05-13 13:05:33 +03:00
Johannes Gäßler	fe0d52b9a2	llama/ggml: add LLM training support (llama/10544) * llama/ggml: add LLM training support more compact progress bar llama_save_model_to_file llama_opt_param_filter ggml_graph_dup force_grads refactor ggml_opt, fix test-opt * remove logits_all * refactor CUDA implementation for ACC * reset graph at beginning of opt period	2025-05-13 13:05:33 +03:00
Dan Johansson	cb90cb0992	ggml-cpu: Integrate fp32=bf16xbf16 SME KleidiAI kernel (llama/13053) * ggml-cpu: Integrate fp32=bf16xbf16 SME KleidiAI kernel Signed-off-by: Dan Johansson <dan.johansson@arm.com> * * code review fixes Signed-off-by: Dan Johansson <dan.johansson@arm.com> * * adds a comment that clarifies barrier usage Signed-off-by: Dan Johansson <dan.johansson@arm.com> --------- Signed-off-by: Dan Johansson <dan.johansson@arm.com> Co-authored-by: Charles Xu <charles.xu@arm.com>	2025-05-13 13:05:33 +03:00
Johannes Gäßler	8264872b5d	CUDA: fix misaligned synchronization in FA (llama/13469)	2025-05-13 13:05:33 +03:00
Atharva Dubey	882d975729	enable dpcpp nightly builds with libraries (llama/13406)	2025-05-13 13:05:33 +03:00
Johannes Gäßler	c426829771	CUDA: fix crash with partial offloading of MoE (llama/13439)	2025-05-13 13:05:33 +03:00
David Huang	0b1962a181	Add `--no-op-offload` to improve `-ot` pp perf in MoE models like llama4 400B (llama/13386)	2025-05-13 13:05:33 +03:00
Johannes Gäßler	86dece9c7c	CUDA: fix race conditions FlashAttention kernels (llama/13438)	2025-05-13 13:05:32 +03:00
Johannes Gäßler	04445664b4	CUDA: fix FlashAttention on Turing (llama/13415)	2025-05-13 13:05:32 +03:00
Jeff Bolz	22f4997dd8	vulkan: scalar flash attention implementation (llama/13324) * vulkan: scalar flash attention implementation * vulkan: always use fp32 for scalar flash attention * vulkan: use vector loads in scalar flash attention shader * vulkan: remove PV matrix, helps with register usage * vulkan: reduce register usage in scalar FA, but perf may be slightly worse * vulkan: load each Q value once. optimize O reduction. more tuning * vulkan: support q4_0/q8_0 KV in scalar FA * CI: increase timeout to accommodate newly-supported tests * vulkan: for scalar FA, select between 1 and 8 rows * vulkan: avoid using Float16 capability in scalar FA	2025-05-13 13:05:32 +03:00
Alberto Cabrera Pérez	b493e03b90	sycl : implementation of reordered Q4_0 MMVQ for Intel GPUs (llama/12858) * sycl : Implemented reorder Q4_0 mmvq Signed-off-by: Alberto Cabrera <alberto.cabrera@codeplay.com> * sycl : Fixed mmvq being called when reorder is disabled * sycl : Improved comments in the quants header Signed-off-by: Alberto Cabrera <alberto.cabrera@codeplay.com> * Use static_assert * safe_div -> ceil_div * Clarify qi comment * change the reorder tensor from init to execute OP * dbg * Undo changes to test-backend-ops * Refactor changes on top of q4_0 reorder fix * Missing Reverts * Refactored opt_for_reorder logic to simplify code path * Explicit inlining and unroll * Renamed mul_mat_algo enum for consistency --------- Signed-off-by: Alberto Cabrera <alberto.cabrera@codeplay.com> Co-authored-by: romain.biessy <romain.biessy@codeplay.com>	2025-05-13 13:05:32 +03:00
Johannes Gäßler	aef59f4851	CUDA: FA support for Deepseek (Ampere or newer) (llama/13306) * CUDA: FA support for Deepseek (Ampere or newer) * do loop unrolling via C++ template	2025-05-13 13:05:32 +03:00
Johannes Gäßler	f8c75dc43e	CUDA: fix crash on large batch size for MoE models (llama/13384)	2025-05-13 13:05:32 +03:00
Radoslav Gerganov	00c8056715	rpc : add rpc_msg_set_tensor_hash_req (llama/13353) * rpc : add rpc_msg_set_tensor_hash_req Use a dedicated struct for the request of RPC_CMD_SET_TENSOR_HASH which makes the code cleaner. * fix	2025-05-13 13:05:32 +03:00
Jeff Bolz	19d8d9a928	vulkan: Allow up to 4096 elements for mul_mat_id row_ids (llama/13326) This assert fired running Qwen_Qwen3-30B-A3B-Q2_K.gguf: GGML_ASSERT(nei0 * nei1 <= 3072); The tensor is 8 x 512. Increase this array size to accommodate.	2025-05-13 13:05:32 +03:00
Alberto Cabrera Pérez	0c4a229154	sycl: addressing non-contiguous src1 mul_mats (nc and batched) (llama/13343) * sycl: fixed non-contiguous src1 mul_mats (nc and batched) * Fixed wrong static_cast inside kernel	2025-05-13 13:05:31 +03:00
R0CKSTAR	09e6b66025	cuda : remove nrows_x in mul_mat_q_process_tile (llama/13325) Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>	2025-05-07 21:00:32 +03:00
Johannes Gäßler	d41cf26a0f	CUDA: mix virt/real CUDA archs for GGML_NATIVE=OFF (llama/13135)	2025-05-07 21:00:32 +03:00
Akarshan Biswas	3c67195be9	SYCL: Disable reorder optimize by default and stop setting tensor extras when optimize is disabled (llama/13254) * SYCL: Do not set tensor extras when reorder optimize is disabled * SYCL: Disable reorder optimize by default	2025-05-07 21:00:32 +03:00
Johannes Gäßler	f9f78a773f	CUDA: fix bad asserts for partial offload (llama/13337)	2025-05-07 21:00:32 +03:00
Johannes Gäßler	be55e25cac	CUDA: fix --split-mode row for MMQ (llama/13323)	2025-05-07 21:00:32 +03:00
Johannes Gäßler	2ffdda99e8	CUDA: fix logic for clearing padding with -ngl 0 (llama/13320)	2025-05-07 21:00:32 +03:00
Akarshan Biswas	9bbedc51cc	SYCL: Disable mul_mat kernels for noncontiguous tensor b (llama/13308) ggml-ci	2025-05-07 21:00:32 +03:00
Diego Devesa	1e1fa27add	rpc : use backend registry, support dl backends (llama/13304)	2025-05-07 21:00:32 +03:00
Aaron Teo	e1bdd148c5	ggml : activate s390x simd for Q3_K (llama/13301) Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>	2025-05-07 21:00:32 +03:00
Johannes Gäßler	7fa8bb303f	CUDA: fix race condition in MMQ stream-k fixup (llama/13299)	2025-05-07 21:00:32 +03:00
Johannes Gäßler	7564f5e6f1	CUDA: fix race condition in MMQ ids_dst (llama/13294)	2025-05-07 21:00:32 +03:00
Jeff Bolz	22ba2e27ce	vulkan: Additional type support for unary, binary, and copy (llama/13266) Support f16->f32 copy. Support f16->f16 and f32->f32 unary ops. Support all combinations of f16/f32 for src0/src1/dst for add/sub/mul/div.	2025-05-07 21:00:32 +03:00
Georgi Gerganov	5eac2a3fbb	vulkan : fix lint (llama/0)	2025-05-07 15:39:32 +03:00
shalinib-ibm	42938398f9	ggml : Enable MMA for BF16 in llamafile_sgemm (llama/13148) This patch upstreams llamafile's cpu matrix multiplication kernels for ppc64le using MMA builtins for BF16 data type. This change results in 9x - 40x gains in total speed S t/s (ie all tokens/total time), across various batch sizes tested using llama-batched-bench benchmark. The patch is tested with Meta-Lllama-3-8B, and Mistral-7B models (BF16 models generated by using llama-quantize from corresponding FP32 models) on an IBM POWER10 machine. Signed-off-by: Shalini Salomi Bodapati <Shalini.Salomi.Bodapati@ibm.com>	2025-05-07 15:39:32 +03:00
Justin Santa Barbara	a8fe90ae15	rpc : avoid uninitialized memory in serialize_tensor (llama/13210) Zero out the name and padding buffers.	2025-05-07 15:39:32 +03:00
Jesse Gross	c5a5a2da5b	ggml: Don't assert fail when tensor data changes (llama/13222) The following scenario will cause an assertion failure in the graph allocator: - Build and allocate a graph containing a tensor with a non-NULL data pointer - Build and allocate a new graph where that data is NULL Result: ggml-alloc.c:819: GGML_ASSERT(talloc->buffer_id >= 0) failed This happens during revalidation because we think that memory should have been previously allocated based on the current graph but in reality the previous graph was different. In this situation, we should do a full reallocation pass.	2025-05-07 15:39:32 +03:00
Diego Devesa	8316bfd82b	build : fix build info on windows (llama/13239) * build : fix build info on windows * fix cuda host compiler msg	2025-05-07 15:39:32 +03:00
Jeff Bolz	fd1cb9fc12	vulkan: Add bfloat16 support (llama/12554) * vulkan: Add bfloat16 support This adds bfloat16 matrix multiply support based on VK_KHR_shader_bfloat16. The extension is required for coopmat multiply support, but matrix-vector multiply trivially promotes bf16 to fp32 and doesn't require the extension. The copy/get_rows shaders also don't require the extension. It's probably possible to fall back to non-coopmat and promote to fp32 when the extension isn't supported, but this change doesn't do that. The coopmat support also requires a glslc that supports the extension, which currently requires a custom build. * vulkan: Support bf16 tensors without the bf16 extension or coopmat support Compile a variant of the scalar mul_mm shader that will promote the bf16 values to float, and use that when either the bf16 extension or the coopmat extensions aren't available. * vulkan: bfloat16 fixes (really works without bfloat16 support now) * vulkan: fix spirv-val failure and reenable -O	2025-05-07 15:39:32 +03:00
Jeff Bolz	17f6b8225e	vulkan: Handle src1 batch dimension in non-contiguous mat-vec-mul shader (llama/13191) * vulkan: Handle src1 batch dimension in non-contiguous mat-vec-mul shader	2025-05-07 15:39:32 +03:00
Acly	6374ea32ca	vulkan : kernels for depthwise 2D convolution (CONV_2D_DW) (ggml/1204) * vulkan : add kernels for depthwise 2d convolution (OP_CONV_2D_DW) * review: remove src_x/y < 0 checks; add performance tests	2025-05-07 15:39:32 +03:00
Daniel Bevenius	09846f4e12	whisper: remove MSVC warnings pragmas (#3090 ) * ggml : remove MSVC warnings pragmas This commit removes the MSVC-specific pragmas as these are now handled in CMakeLists.txt. * whisper : remove MSVC warning pragmas This commit removes the MSVC-specific pragmas. These are now handled in the CMakeLists.txt file.	2025-05-05 13:09:35 +02:00
Jared Tweed	9f540ad8cb	cmake : removed stdc++fs (#3097 ) * removed stdc++fs * kept line, but removed stdc++fs	2025-05-02 12:41:35 +03:00
Johannes Gäßler	d052e64d42	CUDA: batched+noncont MMQ, refactor bs>1 MoE code (llama/13199)	2025-05-01 13:29:02 +03:00
Jeff Bolz	780750a108	vulkan: use uint array index to avoid glslang bug (llama/13193)	2025-05-01 13:29:02 +03:00
shalinib-ibm	919c78e618	ggml : fix ppc64le build (llama/13176) Build fails with compilation error on power pc. This patch fixes the same. Tested with unit tests run via --build <build_dir> && cd <build_dir> && make test Signed-off-by: Shalini Salomi Bodapati <Shalini.Salomi.Bodapati@ibm.com>	2025-05-01 13:29:02 +03:00
Aaron Teo	dc288f84cd	feat(ggml-cpu): enable z17 compile (llama/13182) z17 compilation requires GCC 15.1.0 and onwards Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>	2025-05-01 13:29:02 +03:00
Johannes Gäßler	1543a3600c	CUDA: fix non-cont. inputs for batched mat mul (llama/13155)	2025-05-01 13:29:02 +03:00
Ville Vesilehto	4872355f6e	fix(rpc): Improve input validation and error handling (llama/13069) * fix(rpc): Improve input validation and error handling The `rpc-server` was vulnerable to Denial of Service attacks via several RPC commands (`SET_TENSOR`, `GRAPH_COMPUTE`, etc.). Malformed messages could trigger failed assertions (e.g., invalid `ggml_type`) or out-of-bounds reads/writes leading to `GGML_ABORT` calls, crashing the server process. This PR introduces robust input validation and replaces `abort()` calls with graceful error handling: - Type Validation: `deserialize_tensor` now checks if the `tensor->type` is within the valid `GGML_TYPE_COUNT` range before calling `ggml_new_tensor_4d`. Returns `nullptr` on invalid type. - Bounds Checks: Replaced `GGML_ABORT` in `set_tensor`, `set_tensor_hash`, and `get_tensor` handlers with error logging and returning `false` when data/offset parameters are out of buffer bounds. - Size Checks: Added safe arithmetic checks (for overflow) in `graph_compute` when calculating required message sizes based on client-provided `n_nodes` and `n_tensors`. Returns early if the reported sizes conflict with the actual message size or would lead to overflow. - Error Propagation: - `create_node` now checks for `nullptr` return values from `deserialize_tensor` and its recursive calls, propagating `nullptr` upwards on failure. Uses `find` instead of `at` for safer map access. - `copy_tensor` now checks for `nullptr` from `deserialize_tensor` and sets the response status to failure if deserialization or bounds checks fail. - `graph_compute` now checks for `nullptr` return from `create_node` and returns failure status correctly. The final return value now reflects the actual computation status. These changes improve the RPC server's resilience against malformed client requests, preventing crashes and ensuring errors are handled more gracefully. Signed-off-by: Ville Vesilehto <ville@vesilehto.fi> * refactor(rpc): address pr comments removed comments and unnecessary returns Signed-off-by: Ville Vesilehto <ville@vesilehto.fi> * refactor(rpc): ambiguous nullptr from create_node rpc_server::create_node could previously return nullptr if the input ID was 0 (valid) or if an internal error (deserialization, recursion failure) occurred (invalid). This ambiguity made error handling difficult for the caller (`graph_compute`). This commit clarifies the meaning of nullptr: - `graph_compute` now checks if the input 'id' was non-zero when `create_node` returns nullptr, correctly identifying failures versus intentional null links. - `create_node` avoids recursive calls for zero IDs and propagates nullptr unambiguously on failure during recursion. Signed-off-by: Ville Vesilehto <ville@vesilehto.fi> * refactor(rpc): initial zero check in create_node The caller (`graph_compute`) already checks `id != 0` when handling a `nullptr` return from `create_node`, correctly distinguishing intentional null links from actual errors. This makes the initial `if (id == 0)` check redundant. Also removes the log message when a tensor ID is not found in the provided map which was added in this branch. Signed-off-by: Ville Vesilehto <ville@vesilehto.fi> * fix(rpc): Handle get_alloc_size failure in server Check the return value of `server.get_alloc_size` in the RPC server loop. If the call fails, return early to close the connection. Signed-off-by: Ville Vesilehto <ville@vesilehto.fi> * refactor(rpc): input size validation in graph_compute Removes detailed, step-by-step size calculations and overflow checks in favor of simpler direct comparisons, assuming 64-bit overflow is unlikely. Signed-off-by: Ville Vesilehto <ville@vesilehto.fi> * refactor(rpc): remove extra status code setting Removes the explicit setting of `response.result = GGML_STATUS_FAILED` when `create_node` returns `nullptr` within `graph_compute`. Primary signal is the `false` return value in case of failure. Signed-off-by: Ville Vesilehto <ville@vesilehto.fi> * refactor(rpc): remove redundant check for tensor->type Breaks CI on ubuntu-cpu-make. Tensor type is uint32_t, thus the check is not needed. Signed-off-by: Ville Vesilehto <ville@vesilehto.fi> --------- Signed-off-by: Ville Vesilehto <ville@vesilehto.fi>	2025-05-01 13:29:02 +03:00
Akarshan Biswas	1a76e97c28	SYCL: Add all missing unary kernels (llama/13074) * SYCL: Add all missing unary kernels ggml-ci * decouple kernel launch range from data size using strided loop * use ciel_div helper for num_blocks ggml-ci * clean auto imported header files	2025-05-01 13:29:02 +03:00
R0CKSTAR	7017c1d37d	musa: fix typo in cc control (llama/13144) Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>	2025-05-01 13:29:02 +03:00
Johannes Gäßler	670bf02662	CUDA: fix q_nope_absorbed prec for DS 2 Lite f16 (llama/13137)	2025-05-01 13:29:02 +03:00
R0CKSTAR	9fff2f751c	musa: fix build warning (llama/13129) Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>	2025-05-01 13:29:02 +03:00

1 2 3 4 5 ...

819 Commits