How to turn audio into an SRT subtitle file
An SRT file is plain text: numbered captions, each with a start time, an end time and a line or two of words. Making one by hand from a recording means listening, typing and timing every line. A speech-to-text model does the listening and the timing for you, and what is left is checking the words. This is how to go from an audio or video file to a working SRT, what to watch for, and how to do it without signing up for a monthly plan.
What you need before you start
One audio or video file with the speech you want captioned. CaptionIt reads MP3, M4A, WAV, OGG, FLAC and WebM audio, and MP4, MOV, WebM and MPEG video.
The file can be up to 25 MB and up to 120 minutes long. A long video is usually bigger than 25 MB because of the picture, not the sound. Export the audio track only (most editors call it "export audio" or "share as audio"), or export at a lower bitrate, and the same recording fits easily. Subtitles only need the speech.
From audio to SRT in three steps
1. Open CaptionIt and choose your file under "Audio to SRT". The page reads its length and shows how many minutes it will take before you spend anything.
2. Transcribe the first 30 seconds free, with no account and no card, to judge the quality on your own recording. The free sample takes files up to 3 MB; for a bigger one, export a short clip first.
3. If it is good, sign in and transcribe the whole file. The result lands in the subtitle toolkit on the same page as a timed caption file.
From there you can download it as SRT, convert it to WebVTT for a web player, or take plain text if you only wanted the words.
Check the words, then the timing
Speech recognition gets most words right on clear speech and stumbles on names, jargon and two people talking at once. Read it through once with the video playing, and fix names first, because viewers notice those.
If every caption is early or late by the same amount, shift the whole file. If the captions start in sync and drift further out as the video goes on, the subtitle file and the video disagree about the frame rate; the toolkit fixes that with one multiplier instead of line by line. Long lines can be re-wrapped to a length that reads on a phone.
Paying for one file, not a month
The subtitle toolkit is free. Transcription costs minutes, because running the speech model costs real money. One minute of audio uses one minute of credit, rounded up. Minutes come in one-off packs of 200 or 600 minutes: one purchase, no subscription, and the minutes never expire, so a single interview this month and another next year come out of the same pack.
The audio is transcribed and dropped. Nothing is stored on our side.
Frequently asked questions
Can I convert an MP3 to SRT for free?
The first 30 seconds are free, with no account, as a real caption file you can download (the free sample takes files up to 3 MB). The full recording is paid in minutes from a one-off pack; there is no subscription.
Does it work with video files?
Yes: MP4, MOV, WebM and MPEG. If the video is over 25 MB, export the audio track only and upload that; the captions are the same.
What if my subtitles are out of sync afterwards?
Shift the whole file when every caption is off by the same amount, or apply a frame-rate correction when the gap grows through the video. Both are in the free toolkit on the same page.