How it works

From a recording to a transcript you can trust

The short version is three steps. This is the long version: what happens to your file at each stage, what we can promise and what we can't.

The three steps, in detail

  1. 01

    You drop the file in

    Drag it onto the page or pick it from your computer. It uploads straight to our storage, not through the app, which is why a two-hour recording doesn't time out halfway.

    • The bar shows bytes actually leaving your machine, not a guess.
    • Close the tab if you want: the upload is what matters, and it finishes on its own.
    • Nothing starts processing until the file has landed complete.
  2. 02

    We listen to it

    We measure the real length of the audio, detect the language, separate who says what and write the text with a timestamp on every word.

    • Speakers come out as «Speaker A» and «Speaker B»; you rename them once and the whole transcript follows.
    • The timestamps are per word, not per paragraph, which is what makes the player follow the reading.
    • You see the progress move because it comes from the job itself, not a spinner on a timer.
  3. 03

    You correct and take it away

    Click any line to fix it. The original stays underneath, so a correction is never a loss, and the search index updates with your version.

    • Search inside the text, not just in the titles.
    • Export to DOCX, PDF, SRT and VTT — the subtitle formats keep the timings.
    • Ask for a summary with key points and action items if the recording is a meeting.

Under the hood

What we do that you don't see

Three decisions that explain most of the difference between a usable transcript and a wall of text.

We measure the audio ourselves

The billable minutes come from reading the file, not from asking the transcription provider. If the two ever disagree, you pay for what your file actually lasts.

The raw result is kept

We store exactly what the model returned before we touch it. If we improve how we build paragraphs, we can rebuild your old transcripts without charging you twice.

Speakers live in their own place

A name is stored once, not repeated on every line. That's why renaming someone is instant and two lines can never disagree about who spoke.

Formats and limits

What you can send us

Audio

MP3, WAV, M4A, AAC, FLAC, OGG, OPUS

Video

MP4, MOV, WEBM, MKV, AVI

Maximum size
10 GB per file
Maximum length
10 hours per recording
Languages
90+, detected automatically
Speakers
Up to 32 in the same recording

Video is fine: we only read the audio track, so you're not paying for the picture.

Accuracy

What to expect, honestly

On clean Spanish or English audio we sit around 98% word-for-word. With background noise, people talking over each other or strong accents, the usual range is 92% to 96%.

What makes it better

  • A microphone close to whoever is talking, even a cheap one.
  • People taking turns: overlapping speech is what hurts most.
  • Telling us the language when you already know it, instead of leaving it on automatic.

What makes it worse

  • Phone recordings from across the room.
  • Music or television underneath the conversation.
  • Four people in a café at rush hour.

Your recordings

Where your audio lives

In the European Union

Both the file and the text stay on servers in the EU, and so does the transcription itself.

It gets deleted on its own

Every plan has a retention window. When it's up, the audio and the text go, without you having to remember.

It's yours

We don't train anything with your recordings. Delete a transcript and the audio goes with it, the same day.

Try it with a real recording

120 free minutes a month. No card, and no need to talk to anyone.