Transcribe video to text, without uploading the video

Drop a video onto the page and Whisper writes down its speech in your browser, each line with its start and end time. The video never leaves your machine. Free, with no account.

Free. No account, no watermark. MP4, MOV, WebM and MKV.

How to transcribe a video here

OpenSubs is a subtitle generator that runs in the browser tab, and the first thing it does with a video is transcribe it. The speech is written down on your own machine, as lines of text that each carry a start and an end time.

  1. Open opensubs.app and choose a video, or drop one anywhere on the page. It takes MP4, MOV, WebM and MKV.
  2. Pick a speech model and generate. The model downloads once, when you generate, and is cached by your browser after that.
  3. Check the transcript against the video. The timestamp beside each line is a button: click it and the video jumps to that moment.
  4. Correct any line that came out wrong by typing over it. Every line is a plain text box.
  5. Export the transcript as .srt, .vtt or .ass.

Video transcription that runs on your machine

The usual way to transcribe a video to text is to upload it to a server and wait. Here the work happens in the tab.

Whisper, in your browser

Speech recognition is OpenAI's Whisper, run in your browser, on the GPU where your machine has one. More about Whisper in the browser

The video is not uploaded

There is no upload endpoint in the product. The file is read directly in your browser. Only the model is fetched from the network; the audio is not.

Free, with no account

Transcription and every export work with no account and no sign-in, and they are free. There is no file-size limit, no queue and no watermark.

Two languages in one video

The spoken language is detected every four seconds rather than once per file, so a video that switches between two languages comes out as one file containing both.

The video transcript you leave with

The transcript is exported as .srt, .vtt or .ass. These are subtitle files, and every line in them carries its start and end time. An .srt is a plain text file holding a line number, a start and end time, and the words, so a text editor opens it and anything that takes subtitles will read it.

The same transcript can go further without leaving the page: translate it into another language with the timings left where they were, or burn it into the video as subtitles and export an MP4. If the video is already playing in a tab rather than sitting in a file, the browser extension for Chrome, Edge and Firefox transcribes it as it plays.

What it will not do

Transcription is good and it is not perfect. These are the limits worth knowing before you start.

  • The editor changes words, not timings. You can rewrite what a line says, but not when it starts or how long it stays, and lines cannot be added, deleted or merged.
  • The Tiny model is English only. The three sizes are Tiny (English only, about 40 MB), Base (about 80 MB) and Small (about 250 MB).
  • Accuracy is Whisper's accuracy at the model size the machine can afford, and a browser is a smaller budget than a server.
  • The first visit downloads a model. That is a real wait, and it is the price of not uploading the video.
  • It needs a browser with WebAssembly: Chrome, Edge, Safari or Firefox. WebGPU makes transcription faster but is not required. iOS 17 Safari cannot run the recogniser; iOS 26 works.

What leaves your machine

On the default, nothing. You can check it rather than take it on trust: open the Network tab and transcribe a video. After the page and the speech model have loaded, there are no requests.

Cloud transcription is an option you have to choose. If you choose it, the audio span you trimmed is sent to the service whose key you supplied, and never the video. The on-device option sends nothing and is the default.

The source is public under the AGPL-3.0, on GitHub. Questions, or a problem to report: email hello@opensubs.app.

Questions

How do I transcribe a video to text for free?

Open opensubs.app, drop the video onto the page and generate. Whisper runs on your machine, so nothing is uploaded and nothing is charged. Correct any line by editing its text, then export the transcript as .srt, .vtt or .ass. No account and no watermark.

Is my video uploaded when I transcribe it?

No. There is no upload endpoint in the product. OpenSubs reads the file directly in your browser, and the transcription runs on your own machine. Only the speech model is fetched from the network, never the audio.

Which video formats can I transcribe?

MP4, MOV, WebM and MKV. There is no file-size limit and no queue.

What file does the video transcript come out as?

An .srt, .vtt or .ass file. Every line in it carries its start and end time. An .srt is a plain text file holding a line number, a start and end time, and the words, so a text editor opens it and anything that takes subtitles will read it.

Can I edit the transcript?

You can rewrite the text of any line. The start and end times are not editable, and lines cannot be added, deleted or merged.

Which speech model does the video transcription use?

OpenAI's Whisper, run locally in the browser. Three model sizes are offered: Tiny (English only, about 40 MB), Base (about 80 MB) and Small (about 250 MB). The model downloads once and is then cached by the browser.