Transcribe audio
Turn speech into text with AI (Whisper)
No database: AIMRAN doesn't store your files, text or results
What is Transcribe audio?
This tool turns what is said in an audio or video file into written text: interviews, lectures, meetings, voice notes or videos you need subtitles for.
It uses Whisper, an artificial intelligence speech recognition model. You can choose between a fast mode and a more accurate one, and download the result as text or as a subtitle file with the timing of each sentence.
How to use it
- Upload the audio or video you want to transcribe.
- Choose the model: Fast or Accurate.
- Choose the language: Automatic, Spanish or English.
- Click Transcribe and wait for it to finish.
- Copy or download the text, or download the subtitles as SRT or VTT.
Advantages
- Speech recognition with Whisper, in two sizes depending on whether you prefer speed or accuracy.
- SRT and VTT subtitles with timestamps, ready for video.
- Accepts audio and video in the usual formats.
- Free, no sign-up and unlimited use.
Technical details
Transcription uses the onnx-community Whisper tiny (about 45 MB) and Whisper base (about 80 MB) models, quantized to 8 bits and run with Transformers.js, with WebGPU acceleration when available and CPU (WASM) otherwise. The model is downloaded only the first time. The audio is decoded, mixed down to mono and resampled to 16 kHz, which is the input Whisper expects. The language can be detected automatically or set to Spanish or English. You get the full text plus segments with start and end times, exportable to TXT, SubRip (.srt) and WebVTT (.vtt). Audio shorter than 0.3 s is not processed. Limit of 300 MB per file; long recordings may take several minutes.
Frequently asked questions
How do I convert audio to text for free?
Upload the audio, choose the model and language, and click Transcribe. When it finishes you can copy the text or download it as TXT.
Can I create subtitles for a video?
Yes. Upload the video, transcribe it and download the SRT or VTT file, which includes the timing of each sentence.
What is the difference between the Fast and Accurate models?
Fast uses Whisper tiny: a smaller download and quicker results, with slightly less accuracy. Accurate uses Whisper base: it takes longer but makes fewer mistakes.
Which languages does it work with?
You can set Spanish or English, or leave it on Automatic so the model detects the language.
Why is it slow the first time?
Because the AI model (45 or 80 MB) is downloaded. After that it is already available and starts faster.
More Audio tools
- Convert audio · MP3, WAV, OGG, M4A and FLAC
- Compress audio · Reduce file size by adjusting the bitrate
- Trim audio · With a visual waveform
- Merge audio · Combine several tracks into one
- Volume and normalize · Raise, lower or normalize the level
- Voice recorder · Record from your microphone and download