Meeting notes
Turn a recorded call into text you can search, quote and paste into the minutes. The recording never leaves your laptop.
Transcribed on your device, nothing uploaded.
Drop an audio or video file hereTranscribe an audio or video file
MP3, M4A, WAV, MP4, MOV, WebM and more. Up to 3 hours.
No account, no upload, no waiting in a queue. Speech is recognized on your own device.
The recording is cut at natural pauses and each piece is shown the moment it is done. On a long file you can start reading while the rest is still running.
Your audio is never sent to a server. It is transcribed inside your browser, on this device.
Every sentence carries its time. Click it and the original audio plays from there, so checking a name takes one click.
Drop an MP3, M4A, WAV, MP4 or MOV file, or record straight from the microphone.
Long silences are skipped, and lines like "Thank you." are dropped when nobody is actually speaking underneath.
Correct the text in the editor, then download SRT or VTT. Your corrections keep the original timing.
The first transcription downloads the speech engine once. After that, it's the same three steps every time.
Audio or video, up to 3 hours. Choose the spoken language, or leave it on auto-detect.
Sentences appear line by line. Click any time chip to hear that part.
Edit the text, then download TXT, SRT, VTT or Word, or copy it to the clipboard.
The same tool, set up for different jobs.
Turn a recorded call into text you can search, quote and paste into the minutes. The recording never leaves your laptop.
Transcribe a lecture recording and click any line to replay the part you missed.
Get a first draft of an interview transcript, then check quotes against the audio line by line.
Record a thought on your phone and get it back as text you can paste anywhere.
Most online transcription sites upload your file to their servers and limit free minutes. This one runs on your device.
| speechtotext.tools | Typical online transcription | |
|---|---|---|
| Where your audio goes | Stays on your device | Uploaded to their servers |
| Account | Not needed | Usually required |
| File length | Up to 3 hours per file | Limited on free plans |
| Subtitles (SRT, VTT) | Included | Often on paid plans |
Drop an audio or video file onto the box, or press Record and speak. The sound is converted on your device into a 16 kHz mono track, cut into pieces at natural pauses, and transcribed piece by piece. Each sentence appears as soon as its piece is done, so on a long recording you can start reading while the rest is still being worked on.
The first time you transcribe, the browser downloads the speech engine once. The progress card shows the real size and speed. After that it is stored by the browser and the next file starts right away.
There is no upload step. The file is read and transcribed inside the browser tab, and the text goes straight into the editor. Interviews, patient notes and client calls never leave the device.
Recordings made with the Record button are kept in memory only and are gone when you close the tab. The transcript text is saved in your browser so a reload does not lose it.
Every line starts with a small time chip. Click it and the original audio plays from that sentence, so checking a name or a number takes one click. While the audio plays, the line being spoken is marked in the editor.
Edit the text like any document. The times stay attached to their lines, so the SRT and VTT files you download afterwards carry your corrections with the original timing.
Fast is the default. For English it uses an engine built only for English that works on the actual length of each piece instead of padding every piece to 30 seconds, so short recordings come back quickly. Other languages use a compact multilingual engine.
Accurate uses a larger multilingual engine and a bigger one-time download. In our tests on a laptop, with the processor doing all the work, a 22-second English recording took about 1 second in Fast and 13 seconds in Accurate, and both got the names, dates and the dollar amount right. When a name or an accent keeps coming out wrong in Fast, try Accurate before fixing it by hand.
Speech engines tend to invent text where nobody is speaking. In our tests a 32-second clip of wind and traffic came back as "Thank you." Long silences are therefore never sent to the engine, and short lines such as "Thank you" or "Thanks for watching" are removed only when the audio under them does not look like speech. A real "thank you" at the end of a talk stays.
TXT gives you plain paragraphs: sentences are joined, and a new paragraph starts at pauses of 1.5 seconds or longer. SRT and VTT are subtitle files with at most two lines of 42 characters per caption, ready for YouTube, Premiere, DaVinci Resolve or CapCut. Word (DOCX) opens in Word, Pages and Google Docs.
To hear the text read back, use Read aloud in the download menu. It opens TextToSpeech.tools with the transcript already filled in.
One file can be up to 3 hours long. At 16 kHz mono an hour of audio takes about 115 MB of memory, so very long files on a phone may run out of memory before they finish. If that happens, split the file and transcribe it in parts.
Video files work when the browser can decode their sound track. MP4 with AAC audio, MOV, WebM and MKV with Opus are the most common. A file with no sound track gets a clear message instead of an empty transcript.
No. The file is read and transcribed inside your browser tab. Nothing about the audio or the text is sent to a server.
35 languages, including English, Spanish, Chinese, Hindi, Portuguese, French, German, Japanese, Russian, Korean and Arabic. Choose the language before you start, or pick Auto-detect and the language is recognized from the first 30 seconds of speech.
Up to 3 hours per file. Longer recordings need more memory than most phones give a single tab.
Yes. Download SRT or VTT from the menu next to the download button. Captions are split to at most two lines of 42 characters and keep the corrections you made in the editor.
The speech engine is downloaded once the first time: about 67 MB for English in Fast, about 80 MB for other languages in Fast, and about 250 MB for Accurate. The card above the play bar shows the size and how long is left. Later transcriptions skip this step.
Yes. Press Record, allow microphone access, and press Stop when you are done. The recording is transcribed right away and is not saved anywhere.
Yes, in Safari, Chrome and Firefox. Phones transcribe more slowly than computers, and very long files can run out of memory on a phone.
Drop a file or press Record above. The first transcription sets up the speech engine once.
Back to the toolSister sites that work the same way: in your browser, no upload.