All toolsvideo transcription API SRT

Transcribe Video to Text and SRT

Turn the speech in a video or audio file into timestamped text and a ready-to-burn SRT caption file, with an open model that runs on our own workers rather than a third-party speech API. Run the preset directly on this page or call the same agent-ready endpoint from your product.

LIVE TOOL
Run transcribe video
Drop a video, adjust the defaults if you need to, and download the result here.

or ·

Already have an account? Sign in

Useful defaults, typed options.

The tool slug stays stable while your agent supplies named media inputs and a narrow set of documented options.

  • Self-hosted model: audio never leaves the render worker
  • 25 European languages, detected automatically
  • SRT output drops straight into add-subtitles-to-video
What this preset does
  1. 01Extract the audio track as 16 kHz mono
  2. 02Cut it into utterances with voice activity detection
  3. 03Recognize each utterance with Parakeet-TDT 0.6B v3 (punctuation, capitalization, word timestamps)
  4. 04Return JSON with text and segments plus an SRT caption file
Stable endpointPOST /api/v1/tools/transcribe-video
Live API request
{
  "inputs": [{ "id": "main", "url": "https://example.com/interview.mp4" }],
  "options": { "include_words": false, "speakers": 2 }
}
Successful output example
{
  "result": {
    "output": {
      "filename": "transcribe-video-output.mp4",
      "contentType": "video/mp4",
      "byteSize": 437021,
      "downloadUrl": "https://cdn.kinopipe.com/…"
    }
  }
}

Know the boundaries before you run.

Languages: Bulgarian, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Greek, Hungarian, Italian, Latvian, Lithuanian, Maltese, Polish, Portuguese, Romanian, Russian, Slovak, Slovenian, Spanish, Swedish, Ukrainian. Other languages produce unusable text.
No speaker labels yet; segments follow pauses in speech (up to 20 s each).

About transcribe video

result.analysis has the full text, the speech and audio durations and one segment per utterance with start, end and text (plus per-word times with include_words). result.outputs[1] is captions.srt: cues of at most 6 seconds and two lines of 42 characters, ready for add-subtitles-to-video.

Yes, pass speakers. Give the number of voices when you know it, which is more reliable than making the clustering work it out, or "auto" to let it decide. Each segment then carries a speaker index, the SRT prefixes a turn with "Speaker 1:", and result.analysis.diarization lists the turns with their times. Diarization is a second model on top of the recogniser, so it runs only when you ask.

Then do not use speakers at all. Transcribe each track separately and the labels are exact by construction, with no model deciding anything. Diarization exists for the case where all you have is the mixed file.

NVIDIA Parakeet-TDT 0.6B v3 (CC-BY-4.0) through sherpa-onnx on the same worker that renders your video, after Silero voice activity detection. Nothing is sent to an external speech API.

One credit per started minute of the audio, whatever the worker takes to get through it. A ten-minute interview costs ten credits and a two-hour recording costs a hundred and twenty, so a backlog is priced before you queue it rather than after.

Related media tools