whisper : re-seed decoder 0 between calls (#4025)

Decoder 0 is seeded once in whisper_init_state() and the per-call re-seeding
loop in whisper_full_with_state() starts at j = 1, so decoder 0 is skipped.
Its mt19937 therefore carries over from one call to the next for the whole
lifetime of the state.

It is only consulted in the temperature > 0 branch (whisper_sample_token),
which is reached through temperature fallback -- so the effect only shows on
audio that falls back, and looks like nondeterminism rather than a bug: the
output becomes a function of how many calls the state has already served, and
the same audio decoded twice can yield different text.

Reproduced on a whisper-server instance: three noisy 30 s clips decoded four
times each in the same process gave 3/4 distinct transcripts per clip before
this change and identical transcripts after it. Fresh processes decoding each
clip once hide the issue entirely, which is why it is easy to miss.

Re-seeding decoder 0 with mt19937(0) matches what the loop does for every
other decoder (mt19937(j)) and makes decoding reproducible across calls.

Co-authored-by: david <david@cabrini.ch>
This commit is contained in:
davidcabcab
2026-09-08 11:58:05 +02:00
committed by GitHub
co-authored by david
parent cec4dbe5d5
commit 6d0ed91499
+7
View File
@@ -7006,6 +7006,13 @@ int whisper_full_with_state(
return -4;
}
// decoder 0 is seeded once in whisper_init_state and skipped by the loop below, so its
// generator carries over between calls for the whole lifetime of the state. It is only read
// in the temperature > 0 branch, i.e. on the temperature fallback path, which makes the
// output a function of how many calls the state has already served: the same audio, decoded
// twice, can yield different text. Re-seed it here like every other decoder.
state->decoders[0].rng = std::mt19937(0);
// TAGS: WHISPER_DECODER_INIT
for (int j = 1; j < n_decoders; j++) {
auto & decoder = state->decoders[j];