Audio Transcription (Speech to Text)
Turn speech and audio files into text, locally in your browser with the Whisper model. You pick the language yourself (German or English) and your file never leaves your device.
Your inputs are processed in your browser and are not transmitted to our servers. Note: third-party resources (e.g. advertising and analytics from Google/Cloudflare) and an optional PayPal donation link may transfer data when loading or when clicked. Browser extensions or plugins may be able to read content that is visible in the input fields.
The result will appear here …
Audio Transcription: Turn Speech into Text
Our Audio Transcription tool turns spoken words into text right in your browser. You load an audio or video file, choose the language and get a clean transcript you can copy and reuse in moments. Interviews, meetings, lecture recordings and journalistic material become a written version quickly. Everything runs 100 % locally on your device: your audio is never sent to a server or stored anywhere.
This works through an open-weight speech recognition model that runs directly in your browser. The tool uses the Whisper model in its compact "tiny" variant, running in a Web Worker. This keeps the page responsive even for longer files while recognition works in the background. You choose yourself whether your material is spoken in German or English, and the model works in exactly that language. The model (about 40 MB) downloads on first use and is then cached by your browser.
How it works
Choose an audio file (MP3, WAV, M4A, WebM and more), set the language, either German or English, and start the transcription. The tool recognizes the spoken words and delivers the text directly on the page. You can select and copy the transcript or reuse it for your notes, articles and documents. A subtitle file (SRT) export is not part of this tool yet.
What the tool is good for
- Interviews: turn quotes and conversations into written text after recording
- Meetings: convert calls and video conferences into notes
- Lectures: capture recorded talks and lessons as text
- Journalistic material: review research conversations faster
- Voice memos: bring your own dictations and thoughts into text without typing
What the tool can and cannot do
You get the best results with clear speech and reasonable audio quality. Extremely noisy audio or a very strong accent noticeably reduce accuracy. Very long files take longer and use more device resources. The model is deliberately compact so it can run locally without a server, which is why it works best on simple, clearly spoken language. An SRT export for subtitles is not available yet; the focus is on the plain text transcript.
Note on accuracy: The quality of the transcription depends largely on the audio quality. Clear recordings with distinct speech give the best results; heavy background noise, several overlapping speakers or strong accents can cause errors. So review the transcript before you reuse it. As of September 2026.
Frequently asked questions
Which files can I use?
Common audio formats such as MP3, WAV, M4A and WebM, plus the audio track from video files. For the best accuracy, pick a recording with clear speech and as little background noise as possible.
Are my audio files sent to a server?
No. The transcription runs entirely in your browser. Your audio never leaves your device, it is not transmitted to a server and not stored anywhere.
Do I have to choose a language?
Yes. You choose yourself whether your audio is spoken in German or English. The model then works specifically in that language, which clearly boosts accuracy.
Is the speech recognition model downloaded?
On first use your browser downloads the compact Whisper model (about 40 MB). Afterwards it is cached and available directly for your next transcriptions.
How accurate is the recognition?
With clear speech and reasonable audio quality you get very good results. Heavily noisy audio or a very strong accent can reduce accuracy. For long recordings, allow more time.
Can I export subtitles (SRT)?
Not yet. An SRT subtitle export is currently not part of this tool. The focus is on the plain text transcript that you can copy and reuse.