Run Whisper in your browser, with nothing to install
Whisper is an open-source speech recognition model released by OpenAI. OpenSubs is a subtitle generator that runs it locally in a browser tab. This page is about using Whisper inside OpenSubs. It is not Whisper's official page, and it is not where the model is published.
Free. No account, no API key. Tiny, Base or Small.
How OpenSubs runs Whisper
OpenSubs lives at opensubs.app. Its speech recognition is OpenAI's Whisper, run locally: nothing is installed, and the audio is not sent anywhere.
ONNX weights, in the tab
Whisper is compiled to ONNX and run by transformers.js, inside the browser tab.
WebGPU, or WebAssembly
It runs on WebGPU where the machine has it and on WebAssembly otherwise. WebGPU makes transcription faster but is not required.
Downloaded once
The first time you transcribe, ONNX Whisper weights are downloaded from huggingface.co and cached by your browser. That is a download, not an upload: no audio is in the request.
Largely offline after that
The speech model is cached by your browser after its first download, and everything except cloud translation runs locally.
Nothing to install
It runs in the browser, with nothing installed. There is no Python environment to set up and no command to type.
No account, no API key
Transcription on your device is the default, and it works with no account and no sign-in. A key is only for a cloud service, if you choose one.
Three Whisper model sizes: Tiny, Base and Small
You pick the model. It downloads once and is cached after that.
- Tiny: English only, about 40 MB.
- Base: about 80 MB.
- Small: about 250 MB.
Those are the three sizes offered, and there is no larger one to choose. Accuracy is Whisper's accuracy at the model size the machine can afford, and a browser is a smaller budget than a server. How fast it runs depends on your device.
What you can do with Whisper here
OpenSubs is built around subtitles, so that is the shape the output takes: lines of text with a start and an end time.
Transcribe a video
Drop a video onto the page and its speech is written down on your machine. MP4, MOV, WebM and MKV. Transcribe a video to text
Correct what it got wrong
Every cue is a plain text box. Rewrite the wording and the captions over your video update as you go.
Translate the result
Into 20 languages. Chrome's built-in translation runs on your device at no cost. How subtitle translation works
Burn it into the video
Export a finished MP4 with the subtitles in the pixels, encoded in the browser with WebCodecs. Burn subtitles into a video
Export a subtitle file
Take the .srt, .vtt or .ass and use it elsewhere.
The extension and the desktop app
The page needs nothing installed. These two are installed, and they do things a page cannot.
Browser extension
It transcribes whatever plays in your tab as it plays, using a local Whisper model: Tiny, Base or Small, on WebGPU or CPU. The subtitles extension for Chrome, Edge and Firefox
Desktop app
For macOS, Windows and Linux, and the one for long files and batches. With its local features, speech recognition takes place on your device. Downloads
What to expect
Running the model on your own machine has a cost, and it is better to hear it here.
- The first visit downloads a model. That is a real wait, and it is the price of not uploading the video.
- The Tiny model is English only.
- It needs a browser with WebAssembly: Chrome, Edge, Safari or Firefox. iOS 17 Safari cannot run the recogniser; iOS 26 works.
- The editor changes words, not timings. Cue text is editable; the start and end times are not.
OpenSubs is open source under the AGPL-3.0, so you can read what runs on your machine on GitHub. Questions, or a problem to report: email hello@opensubs.app.
Questions
Is this the official Whisper page?
No. Whisper is an open-source speech recognition model released by OpenAI. OpenSubs is a subtitle generator that runs Whisper locally in your browser, and this page describes using Whisper inside OpenSubs.
Do I need to install anything, or have an API key, to use Whisper here?
No. OpenSubs runs in the browser at opensubs.app with nothing installed. Transcription on your device is the default and works with no account, no sign-in and no API key. The speech model is downloaded once and then cached by the browser.
Which Whisper models can I choose?
Three sizes are offered: Tiny (English only, about 40 MB), Base (about 80 MB) and Small (about 250 MB). There is no larger size to choose.
Is my audio or video uploaded?
No. Only the model is fetched from the network, never the audio. There is no upload endpoint in the product, and the video is read directly in your browser.
Does Whisper use my GPU in the browser?
It runs on WebGPU where the machine has it and on WebAssembly otherwise. WebGPU makes transcription faster but is not required, so how fast it runs depends on your device.
Does it work offline?
After the first visit, largely yes. The speech model is cached by your browser after its first download, and everything except cloud translation runs locally.