Transcribe an MP3 file
The format everything records in, turned into text with timestamps and speakers. Drop the file and read it in minutes.
- Any bitrate, mono or stereo, any length
- Language detected on its own, 90+ of them
- Text, subtitles and summary from one pass
How it works
- 01
Drop the MP3
Up to 10 GB — hours of audio in a single file.
- 02
It transcribes itself
About three minutes per hour of recording, with progress on screen.
- 03
Download what you need
DOCX, PDF, SRT, CSV or JSON, with your corrections in.
What goes in
- MP3
Up to 10 GB per file, as many files as you want in one batch. The language is detected on its own.
What comes out
- DOCX
- SRT
- VTT
- CSV
- JSON
The same transcript, in whichever file you need. Your corrections are included in all of them.
- Speakers separated
- 90+ languages
- EU servers
The format everything records in, turned into text with timestamps and speakers. Drop the file and read it in minutes.
In detail
MP3 is what a voice recorder, a phone app or a podcast host gives you, and it's a perfectly good format to transcribe from. Compression throws away what the ear can't hear, and speech recognition doesn't need what the ear can't hear.
Bitrate barely matters either. From 64 kbps upwards the accuracy is essentially the same; what does move it is how close the microphone was and whether two people talk at once. A 32 kbps recording from across a room is the one case worth re-recording.
- Any bitrate from 32 kbps up, mono or stereo, constant or variable.
- Long files handled whole: a three-hour podcast doesn't need splitting.
- The ID3 tags don't matter, and neither does the file name.
- Speakers separated even when the MP3 is a mono mixdown of several mics.
What people use it for
Voice recorders
Almost every handheld recorder writes MP3 by default. Drop it in as it comes.
Podcast episodes
The published MP3 is enough: no need to go back for the session files.
Voice notes
A phone memo of an idea at two in the morning is still a transcribable file.
Questions
Start with your own recording
The format everything records in, turned into text with timestamps and speakers. Drop the file and read it in minutes.