Automatic language detection
Let the speech model detect the language, or provide a language hint for common languages.
Turn spoken audio into readable text and timed SRT or WebVTT captions.
How it works
Upload one supported audio file, let AudioTidy transcribe the speech, then save the result you need.
01
Select an AAC, FLAC, M4A, MP3, OGG, WAV, or WebM file up to 20 MB.
02
Choose automatic language detection or provide a language hint, then securely upload the file for processing.
03
Review and copy the text, or download the result as TXT, SRT, or timed WebVTT captions.
Transcription features
A focused transcription workflow with clear status, short retention, and useful output formats.
Let the speech model detect the language, or provide a language hint for common languages.
Use plain text for notes and documents or timed SRT and WebVTT captions for supported media players.
Source audio is deleted after processing and downloadable results expire after 24 hours.
Audio to text FAQ
Learn about uploads, supported formats, languages, output files, and retention.
Yes. After you start transcription, the selected file is securely uploaded to Cloudflare R2 and processed with Cloudflare Workers AI. Editing tools elsewhere continue to process files locally.
The source audio is deleted after transcription succeeds or permanently fails. The transcription result remains available for up to 24 hours for TXT, SRT, and WebVTT downloads.
Audio to Text accepts AAC, FLAC, M4A, MP3, OGG, WAV, and WebM files up to 20 MB.
Automatic detection is the default. You can also hint English, Chinese, Spanish, French, German, Japanese, Korean, Portuguese, Italian, Russian, Arabic, or Hindi.
Download a UTF-8 TXT file, an SRT file, or a WebVTT file containing timed captions.
Yes. Once upload completes, this browser remembers and restores the current task after a refresh.