Extract Audio Timing
Extract speech, audio and word timestamps for captions or video reference analysis.
Read as Markdown ↗Connect one ready Video or Audio to Extract Audio Timing, then choose Extract speech timing. Video audio is extracted in your browser; keep the page open until extraction finishes. The speech service receives only audio. Audio can be up to ten minutes and 20 MB. Longer audio is rejected, never clipped.
The node stays compact and shows when its audio and speech timing are ready. Click View result to open the audio preview, transcript and timing JSON in a dialog. Motion Video's connected reference has a View JSON dialog too. It uses your daily LLM Gem allowance, with a minimum of 1 Gem per extraction. Longer audio can use more Gems based on its speech recognition cost. Successful results are reused for the same audio, including after refreshing. A failed request requires Retry extraction; a timed-out request may have incurred usage.
Connect the single Audio, transcript + timing output to Motion Video for captions and word highlighting. You can also connect the original video: Motion then uses its embedded speech once, with the measured word timing. Connect the same output to a Text node to summarize or reuse the spoken content in another workflow. Timings are seconds from the start of the media, including any opening silence. Changing the source requires extraction again; the previous timing is not reused.
For a Video model that accepts audio references, the same output supplies the prepared audio without requiring the destination to consume the transcript. If you only need sound from a selected video section, use Extract Video Segment, choose Audio only, and connect its audio output directly. This avoids a speech-timing request and its allowance use. Use Smart Extractor when you need the full source audio.
Save this wiring as a group preset to reuse the workflow with new media. Presets retain the wiring and clear the private audio and transcript results.