No media upload
The selected media is processed on your device and is not sent to Whisper Web servers.
Select an audio or video file to create editable text, timestamps or subtitles. Whisper processes the media in this browser. Files can be up to 300 MB and 20 minutes.
No account · No media upload · TXT, SRT, VTT and JSONProcessing happens on this device. Direct URLs must allow browser access (CORS).
Choose a scenario for task-specific steps, or read a guide before deciding where the audio should be processed and how the result should be exported.
Whisper Web decodes media and runs the Whisper model in your browser. Completed transcripts stay in IndexedDB until you export or delete them.
The selected media is processed on your device and is not sent to Whisper Web servers.
WebGPU can run faster; WebAssembly provides a broadly compatible fallback.
Correct the text, copy it or export TXT, JSON, SRT or VTT.
No. In local mode, the browser reads and decodes it in memory.
The selected Whisper model downloads once and is then cached by your browser.
Whisper Web accepts common MP3, WAV, M4A, MP4 and WebM files, subject to your browser's decoding support.
The Audio language menu includes 99 languages, including English, Spanish, Arabic, Chinese, Hindi, French, German, Japanese and Portuguese. Choose the language spoken in the recording before you start.
Use the large-file page for a local audio or video file up to 1 GB and 1 hour. Keep this tab open: a long job can warm the device and take time. A reference estimate appears after processing begins.
Transcribe a large file