Bulk audio transcription API: transcribe a list of files
Send a list of audio or video links (MP3, MP4, podcast RSS feeds, YouTube, Google Drive or Dropbox) to the Audio Transcriber on Apify and get each file back as text with timestamps plus SRT, VTT, TXT and Markdown files, for $0.006 per audio minute (Whisper large-v3-turbo, 90+ languages); failed or silent files are free.
Last updated
How to do it in 4 steps
- Paste the file links into Audio or video file URLs (up to 1,000 files per run).
- Optional: set the Language, turn on Speaker labels ($0.015 per minute instead of $0.006), or cap the cost with Max minutes per file and Max total minutes per run.
- Click Start, or start the same run with one HTTP request:
POST https://api.apify.com/v2/acts/tidytools~audio-transcriber/runswith your Apify token and the same input as JSON. - Read one row per file with
text, timestampedsegments,srtUrlandvttUrl. Rows of failed or skipped files say why and showcharged: false.
Specifics
| Input | Audio and video file links, podcast RSS feeds, YouTube, Loom and Wistia links, public Google Drive, Dropbox and OneDrive share links |
|---|---|
| Output | Text, timestamped segments, paragraphs, SRT, VTT, TXT and Markdown files per file |
| Price | $0.006 per audio minute ($0.36 per hour); speaker labels $0.015 per minute; optional AI summary $0.01 per file; no start fee; billed per started minute |
| Not charged | Files that cannot be downloaded or decoded, files with no speech, files skipped by your limits |
| Limits | Up to 1,000 files per run, up to 2 GB per file, any length |
| Measured speed (our tests) | 8-minute MP4 in 33 seconds, 26-minute MP3 in 109 seconds |
FAQ
Is there a 25 MB file limit?
No. Send the file link instead of the file: files up to 2 GB and any length work, with no OpenAI key.
Which speech-to-text model does it use?
OpenAI's open Whisper large-v3-turbo model, with automatic language detection and 90+ languages. Speaker labels, when turned on, come from nova-3.
How is a file billed?
Per started minute of audio: a 500-second file is billed 9 minutes. SRT and VTT subtitles, paragraphs and timestamps are included.
Can I run it from n8n, Make or Zapier?
Yes. Every run can be started from the Apify API, a schedule, a webhook or the Apify apps for Make, n8n and Zapier.
Related guides
Related
Prices, limits and timings from the Audio Transcriber README on Apify. Prices are the base (Free plan) prices on Apify; larger Apify plans pay less. Actor page: apify.com/tidytools/audio-transcriber.