ggml-quants : avoid invalid rounding in qkx3 scale search (llama/29817)

* ggml-quants : avoid invalid rounding in qkx3 scale search

The imatrix scale search can produce an infinite, NaN, or otherwise out-of-range value when the fitted minimum collapses to the maximum or makes the range extremely small. That value is then passed to nearest_int and can trip its assertion in Debug builds.

Clamp the quantization level to [0, nmax] before rounding so valid in-range values behave the same as before while invalid scale-search results no longer reach nearest_int.

Add regression coverage for degenerate imatrix groups across q2_K, q4_K, q5_K, q4_1, and q5_1.

Fixes #29804.

Assisted-by: Claude Opus 5.5

* tests: print degenerate imatrix quant types
This commit is contained in:
Yash Raj Pandey
2026-10-06 10:38:18 +03:00
committed by Georgi Gerganov
parent 0b35d1f30b
commit 6b704713cc
+3 -2
View File
@@ -1036,8 +1036,9 @@ static float make_qkx3_quants(int n, int nmax, const float * GGML_RESTRICT x, co
iscale = (rmin + rdelta*is + nmax)/(max - min);
float sum_l = 0, sum_l2 = 0, sum_xl = 0;
for (int i = 0; i < n; ++i) {
int l = nearest_int(iscale*(x[i] - min));
l = MAX(0, MIN(nmax, l));
// min is the best fit so far and can be at or near max, so v can be inf, nan or out of range for nearest_int
const float v = iscale*(x[i] - min);
const int l = v > 0 ? nearest_int(MIN(v, nmax)) : 0;
Laux[i] = l;
float w = weights ? weights[i] : x[i]*x[i];
sum_l += w*l;