whisper : expose internal VAD speech segments (#3916)

When transcribing with params.vad = true, whisper already computes the speech
segments and keeps them in the state. Expose them so callers can reuse those
boundaries (for example to align or clip subtitles to speech) instead of running
a second, separate VAD pass.

Times are on the original audio timeline in centiseconds; the count is 0 when VAD
was not used. test-vad-full.cpp checks the segments are ordered and non-empty.
This commit is contained in:
Lin Xiaodong
2026-07-01 13:10:16 +02:00
committed by GitHub
parent 167d225f3a
commit 6fc7c33b4c
3 changed files with 47 additions and 0 deletions
+12
View File
@@ -62,6 +62,18 @@ int main() {
prev_t0 = t0;
}
// internal VAD speech segments, on the original audio timeline (centiseconds)
const int n_vad = whisper_full_n_vad_segments(wctx);
assert(n_vad > 0);
int64_t vad_prev_end = -1;
for (int i = 0; i < n_vad; ++i) {
const int64_t t0 = whisper_full_get_vad_segment_t0(wctx, i);
const int64_t t1 = whisper_full_get_vad_segment_t1(wctx, i);
assert(t1 > t0);
assert(t0 >= vad_prev_end); // segments are ordered and non-overlapping
vad_prev_end = t1;
}
whisper_free(wctx);
return 0;