Skip to content

100% local — your file never leaves your browser

Video Transcript Generator

Paste a link or upload a file. Get a clean transcript in seconds — free, no sign-up, SRT and VTT included.

  • Free
  • No sign-up
  • No watermark
100% local — your file never leaves your browserReady

How it works

  1. 01

    Paste or upload

    Drop in a video link, or pick an audio or video file from your device.

  2. 02

    We transcribe it

    Speech is detected automatically in 90+ languages and turned into text.

  3. 03

    Copy or export

    Copy the transcript, or download it as TXT, SRT or VTT for subtitles.

Good to know

Two modes, and when to use each

Upload mode runs the speech model inside your own browser. Nothing is sent to a server, there is no limit on how much you transcribe, and it keeps working offline once the model has cached. It is the right choice for anything private and for everyday use. Link mode fetches a public post for you, which saves a step when you are working from something already published. The tradeoff is that it depends on the platform cooperating — private posts and region-locked videos cannot be fetched by anyone.

What a transcript is actually good for

Most people do not want the transcript itself. They want what it unlocks: subtitles for a video that will otherwise be watched on mute, a caption written from what was actually said rather than from memory, timestamps that point to the twenty seconds worth clipping, or a searchable record of a meeting nobody took notes in. Text also translates cleanly, which audio does not — a transcript is usually the first step to reaching an audience in another language.

Accuracy, honestly

Clean single-speaker audio comes out close to verbatim, and editing the result is faster than typing from scratch. Accuracy degrades in predictable ways: two people talking over each other blur into one stream, proper nouns and technical jargon are the most common errors, and distant or noisy recordings suffer badly. The model guesses a plausible word rather than leaving a gap, so mistakes read confidently. Treat the output as a strong first draft, not a finished record.

Why it costs nothing

Transcription services normally charge by the minute because every minute you send costs them GPU time. Upload mode moves that work onto the device you are already holding, so there is no per-minute cost to pass on and no incentive to meter you. The model downloads once, roughly the size of a short podcast episode, and is reused from cache from then on. Link mode does involve a server, which is why it carries a daily limit rather than being unlimited.

What we do not keep

In upload mode there is nothing to keep — the file is read from disk by your own browser and never reaches us. In link mode the audio is fetched, transcribed and discarded within the request; the transcript is returned to you and not stored. There is no account, so there is no history to mine and nothing to leak. You can confirm the upload claim yourself in the browser network panel, which is a stronger guarantee than any privacy policy.

Frequently asked questions

Is this video transcript generator free?
Yes. Link transcription is free every day with a generous daily limit, and file upload mode runs entirely in your browser with no limit at all.
Do I need to create an account?
No. There is no sign-up, no email and no watermark on anything you export.
Which languages are supported?
Over 90 languages are detected automatically, including English, Portuguese, Spanish, Vietnamese and Russian. Detection uses the opening of the recording, so a clip that switches language partway through may be labelled by whichever language dominates the start.
Can I get subtitles from the transcript?
Yes. Every transcript can be exported as SRT or VTT with timestamps, ready to drop into any video editor. SRT suits editors and social platforms; VTT is required for HTML5 video on your own site.
How long can the audio be?
There is no server-side limit in upload mode because there is no server. Very long recordings are limited by your device memory — multi-hour files are comfortable on a desktop and can struggle on a phone.
Does it label who is speaking?
No. Speaker labelling (diarization) needs a second model that is not practical to run in a browser. The output is a single continuous text stream.