Describe a Video as Text
Turn a video into a summary and a list of timestamped scene descriptions, so an agent can decide what to trim, split or caption without watching the footage or sending frames to a model on every call. Run the preset directly on this page or call the same agent-ready endpoint from your product.
Useful defaults, typed options.
The tool slug stays stable while your agent supplies named media inputs and a narrow set of documented options.
- One call replaces frame-by-frame vision passes
- Timestamps line up with trim-video and split-video
- Optional focus for what to look for
- 01Detect shot changes and build the scene list
- 02Extract one frame per scene, or one every 8 seconds when there are no cuts
- 03Describe the frames with a vision model in the language you choose
- 04Return a summary and timestamped scenes as JSON
POST /api/v1/tools/describe-videoKnow the boundaries before you run.
About describe video
Related media tools
Transcribe Video to Text and SRT
Turn the speech in a video or audio file into timestamped text and a ready-to-burn SRT caption file, with an open model that runs on our own workers rather than a third-party speech API.
ExploreDetect Scene Changes in a Video
List the timestamps where the picture cuts to a new scene, then pick trim, split and clip boundaries from real cuts.
ExploreSplit Video by Scene Changes
Detect every cut in a video and return one independently playable MP4 clip per scene. Detection and splitting run in a single job.
Explore