Private browser transcription
Audio decoding and Whisper inference run on this device instead of uploading the recording to a transcription server.
Transcribe MP3, WAV, M4A, FLAC, AAC, OGG, Opus, or WebM speech into editable text and SRT subtitles without uploading audio.
Drop speech audio here, or or import from:
An audio to text converter recognizes spoken words in a recording and turns them into editable text. Use it for interviews, podcasts, meetings, lectures, voice memos, and other permitted recordings.
Upload an audio file or record from your microphone, choose the spoken language or use automatic detection, and export a transcript with timestamps. The multilingual Whisper Tiny model runs directly in your browser.
Audio decoding and Whisper inference run on this device instead of uploading the recording to a transcription server.
Use automatic language detection or select English, Chinese, Spanish, French, German, Japanese, Korean, Portuguese, or Italian.
Review speech in timed sections and download SRT subtitles as well as a plain TXT transcript.
Transcribe an existing audio file or create a new microphone recording without installing desktop software.
The first transcription downloads Whisper Tiny and caches it in the browser. Later sessions can reuse that model cache.
Choose an MP3, WAV, M4A, AAC, FLAC, OGG, Opus, or WebM audio file, or record directly from your microphone.
Leave Auto detect selected for mixed workflows or specify the language to improve recognition speed and consistency.
Start Whisper Tiny, review and edit the result, then copy the text or download TXT and timestamped SRT files.
Turn recorded speech into reusable text for editing, publishing, accessibility, and review.
Create a searchable draft from a recorded conversation, research interview, meeting, or voice memo.
Download an SRT file with timed speech sections for tutorials, presentations, and social videos.
Turn spoken episodes into notes, outlines, quotes, summaries, or an editable publishing transcript.
Convert clear lecture or lesson audio into text that can be reviewed alongside the original recording.
Whisper Tiny favors speed and a smaller download. Recording quality still has a major effect on the transcript.
Reduce background music, echo, overlapping speakers, and distant microphone sound whenever possible.
Choose a specific language when automatic detection is uncertain or the recording is short.
Small speech recognition models can miss uncommon names, brands, abbreviations, and specialist vocabulary.
Long audio is processed in 30-second sections. Avoid closing the page until the final transcript appears.
Local and cloud transcription make different tradeoffs between privacy, setup time, speed, and model size.
| Feature | This Whisper Tiny tool | Cloud transcription |
|---|---|---|
| Audio processing | On this device | Uploaded to a remote server |
| First use | Model download required | Usually immediate |
| Offline reuse | Possible after browser caching | Internet connection required |
| Accuracy tier | Fast compact Whisper model | May use larger server models |
| Exports | Editable text, TXT, and SRT | Depends on the provider |
Desktop Chrome, Edge, or Firefox is recommended. Processing time depends on recording length and device performance.
Yes. You can upload or record audio, transcribe it with Whisper Tiny, edit the result, and download TXT or SRT without creating an account.
No. The recording is decoded and transcribed in your browser. Audio and transcript stay on this device.
The browser downloads a quantized Whisper Tiny model on first use. Typical connections may take about 10 to 40 seconds, and the cached model can be reused later.
The uploader accepts MP3, WAV, M4A, AAC, FLAC, OGG, Opus, and WebM. Actual decoding support can vary slightly by browser and operating system.
Yes. Whisper Tiny is multilingual. Use automatic language detection or select one of the provided language options before transcription.
Yes. The timestamped sections can be downloaded as an SRT subtitle file, while the complete editable transcript can be downloaded as TXT.
Whisper Tiny is optimized for compact browser inference. Background noise, overlapping speakers, music, accents, and uncommon names can reduce accuracy, so review the editable transcript before publishing.
This browser implementation accepts recordings up to 60 minutes and 500 MB. Long recordings need more memory and processing time.