Transcribe Video to Text and SRT
Turn the speech in a video or audio file into timestamped text and a ready-to-burn SRT caption file, with an open model that runs on our own workers rather than a third-party speech API. Run the preset directly on this page or call the same agent-ready endpoint from your product.
Useful defaults, typed options.
The tool slug stays stable while your agent supplies named media inputs and a narrow set of documented options.
- Self-hosted model: audio never leaves the render worker
- 25 European languages, detected automatically
- SRT output drops straight into add-subtitles-to-video
- 01Extract the audio track as 16 kHz mono
- 02Cut it into utterances with voice activity detection
- 03Recognize each utterance with Parakeet-TDT 0.6B v3 (punctuation, capitalization, word timestamps)
- 04Return JSON with text and segments plus an SRT caption file
POST /api/v1/tools/transcribe-videoKnow the boundaries before you run.
About transcribe video
Related media tools
Add Subtitles to Video
Burn an SRT or WebVTT caption track into a durable video output.
ExploreDescribe a Video as Text
Turn a video into a summary and a list of timestamped scene descriptions, so an agent can decide what to trim, split or caption without watching the footage or sending frames to a model on every call.
ExploreDetect Silence in Audio or Video
List silent passages with start, end and duration for chaptering, cut planning or ad-break placement, without editing anything.
Explore