Whisper with word-level timestamps and speaker diarization for real subtitles.
Catalog snapshot
BSD-2-Clause; diarization models have their own weights and terms.
Fetched 1
whisperX layers forced alignment and speaker diarization on faster-whisper: word-accurate timestamps and who-said-what. It is the standard for subtitle files from local audio. BSD-2-Clause licensed, permissive like MIT in practice.
word-level timestamps · speaker diarization · faster-whisper core · alignment models