Turn speech into text
Audio and Video Transcriber
Choose a recording, download a Whisper model once, and transcribe it on your device. Edit the result and export TXT, SRT, WebVTT, or timestamped JSON.
Private by default
The recording stays in your browser. The tool downloads a speech model, then performs transcription on your device without an inference API.
Use the tool
Add your file or text below, adjust the options if needed, then download the result.
You can add
MP3, WAV, M4A or AAC, MP4, WebM, OGG, Browser-readable media
You will get
Editable transcript, TXT, SRT, WebVTT, Timestamped JSON
How to use Audio and Video Transcriber
- 1
Choose a recording
Select browser-readable audio or video from your device. Unknown durations are verified during local decoding.
- 2
Run Whisper locally
Download the selected model once, then transcribe the recording on your device.
- 3
Review and export
Correct the editable transcript and download TXT, SRT, WebVTT, or timestamped JSON.
When this tool is useful
- Create a first-pass transcript from a recorded interview or meeting.
- Generate subtitle files for a presentation, tutorial, or social video.
- Turn a voice memo or podcast recording into editable text without uploading it.
Browser and format limits
- The first run downloads a local speech model, and the browser may later remove cached model files.
- Files are limited to 150 MB and recordings are capped at 20 minutes on typical desktops or 10 minutes on smaller and lower-memory devices.
- Transcription speed depends on device memory, browser support, and available acceleration.
- Names, overlapping speakers, music, background noise, and specialized vocabulary can reduce accuracy.
Practical tips
Use clear audio with one speaker at a time whenever possible.
Keep this tab open and prevent the device from sleeping during long transcriptions.
Review names, numbers, and technical terms before publishing the transcript.
Audio and Video Transcriber questions
Is my audio or video uploaded?
No. The media is decoded and transcribed in your browser. The tool downloads model files, but it does not send your recording to AI Guys or an inference API.
Why does the first transcription need a large download?
Speech recognition runs from a Whisper model on your device. Your browser caches the downloaded model when storage is available, so later sessions can reuse it.
Which recordings work best, and how long can they be?
Clear speech with limited background noise works best. The tool supports up to 20 minutes on typical desktops and up to 10 minutes on smaller or lower-memory devices. Codec support depends on your browser.
Related tools
View the full libraryTranscript Cleaner
Remove timestamps, cue numbers, extra spacing, and optional filler words.
Move between caption formatsSubtitle Converter
Convert SRT and WebVTT captions or extract their plain spoken text.
Inspect audio and videoMedia Metadata Viewer
Inspect duration, size, type, and browser-readable audio or video details.
Put file work inside your workflow.
AI Guys builds custom applications and automations when a browser utility needs approvals, integrations, or reliable scale.