← Back to AI Tools

WhisperX Word-Level Timestamp Transcription

WhisperX fills two gaps in OpenAI Whisper: the official repository reports about 70x realtime with large-v2 using VAD segmentation and batching, then aligns output with an external phoneme model to sharpen utterance-level timestamps down to each word, with speaker diarization via pyannote

Tool Interface

Interactive tool will be available soon

Features

  • ✓ Fast: the official repository cites roughly 70x realtime transcription with large-v2, achieved through batched inference and a faster backend
  • ✓ Word-level timestamps: WhisperX aligns output with an external phoneme model for accurate per-word timing, fixing Whisper's utterance-level timestamps that can be off by several seconds
  • ✓ Speaker diarization: available through pyannote, so multi-speaker conversations can be attributed to a specific speaker
  • ✓ VAD preprocessing: voice activity detection segments the audio before recognition, giving cleaner long-form chunks that are easier to batch
  • ✓ Open source and self-hostable: released under a BSD-2-Clause licence, so you can run it on your own hardware instead of shipping audio to a third-party API

How to Use

  1. Read the README and installation notes in the official repository at https://github.com/m-bain/whisperX
  2. Prepare the environment as documented (typically Python and PyTorch, with ffmpeg installed so audio can be decoded)
  3. Transcribe with the documented CLI or example script, adding the diarization options when you need speakers separated (pyannote requires a Hugging Face access token)
  4. Take the word-timestamped, speaker-labelled output into your own subtitles, meeting notes or editing pipeline

FAQ

What is WhisperX?

WhisperX is an open-source transcription tool that builds on OpenAI Whisper; the official repository is https://github.com/m-bain/whisperX . It provides fast automatic speech recognition with word-level timestamps and speaker diarization, and the repository cites around 70x realtime with large-v2.

How is it different from plain Whisper?

Per the repository, Whisper transcribes accurately but its timestamps are utterance-level and can be inaccurate by several seconds, and it does not support batching natively. WhisperX adds batching for speed and aligns Whisper output with an external phoneme/CTC model to get accurate word-level timestamps, plus speaker diarization.

What are word-level timestamps good for?

They give every word a timeline, which matters for subtitles, karaoke-style text highlighting, cutting audio by sentence, jumping from a meeting note to the exact second in the recording, and aligning content per speaker. Compared with utterance-level timestamps, word timing greatly reduces subtitle drift.

What extra setup does diarization need?

The documentation states diarization is available through pyannote and typically requires a Hugging Face access token to pull the diarization models. So the feature works, but it needs an extra step for token and model authorisation.

Can I use it commercially?

The WhisperX code itself is released under the permissive BSD-2-Clause licence, but the models and services it depends on carry their own licences and terms (for example the Whisper models, the pyannote models and Hugging Face terms). Check each one against its official page before commercial use.

What hardware does it need?

The repository notes under 8GB of GPU memory for large-v2 at beam_size 5, so a mainstream NVIDIA card is enough; smaller models or CPU inference also work but run noticeably slower. Actual runtime and memory depend on audio length, model size and hardware, so check the official docs and your own benchmarks.