Turn a recorded call or meeting into minutes you can edit and share.
Your browser downloads the speech model once, then runs it on your own machine. Your audio itself is never transmitted, on the first run or any run. Compare that with services that hold your recording for thirty days.
Every line is timed and editable. Export as text, SRT, VTT, JSON, CSV or Markdown once you have it the way you want.
No sign-up, no email, no monthly minute cap and no paid tier holding the useful part hostage.
Transcribe a Meeting runs entirely inside your browser, so your file is never uploaded anywhere. Speech recognition uses OpenAI's Whisper, the same model most paid services run, except here it runs on your machine rather than theirs. The model is downloaded once, about 76 MB, and cached by your browser. Your recording itself is never transmitted.
Yes, and there is nothing unpleasant to find later. No account, no email, no minute cap, no watermark, and no paid tier holding the useful part. Your own computer does the work, so there is no per minute cost for us to pass on to you.
The first time, your browser downloads a 76 MB speech model. That happens once and is then cached. After that it depends on your device: on the machine this was built on, eight seconds of audio took about eleven seconds with the model already loaded, using the graphics card. A device without WebGPU falls back to the processor and takes longer.
We never receive it. That is worth being exact about, because most transcription services upload your audio to a server and promise to delete it afterwards. Here the audio never leaves your device, so there is no copy to delete, nothing to leak, and nothing anyone could ask us to hand over.
Open your browser's developer tools and watch the network panel. On the first run you will see the speech model being fetched, once. After that you will see nothing at all while it transcribes, because your audio is never sent anywhere.
It uses Whisper, the same model family the paid services use. Clear speech from one or two people comes out very well. Heavy accents, background noise and people talking over each other are where every speech model struggles, and this runs a smaller version than a server would, so it will make more mistakes on difficult audio. Every line is editable before you export.
Video or audio: MP4, MOV, WEBM, MKV, AVI, MP3, M4A, WAV, FLAC, OGG, AAC, WMA and more. The soundtrack is pulled out of a video automatically, so an MP4 behaves exactly like an MP3.
Yes. Every line carries a start and end time, so you can export straight to SRT or WebVTT for a video editor, or to TXT, JSON, CSV or Markdown if you only want the words.
Long video is the current weak point. The whole file is decoded in memory, so a two hour recording can exhaust what a browser allows and fail. Calls, interviews, podcasts and ordinary recordings are fine. Chunked handling for long files is being built.
Once the model has been downloaded, yes. Disconnect from the internet and it keeps transcribing, because everything it needs is already sitting in your browser.
Yes. Record from your microphone, or capture the audio of another browser tab to get both sides of a call. Both can run at once and are mixed into one recording.
Every tool is free, with no account and no limit.