Turns a video into a timestamped, action-by-action written description of what happens in it.
vidannotate run gameplay.mp4
# -> runs/gameplay-1a2b3c4d/final.txt (also final.json, final.srt)New here? GUIDE.md walks through a single video end to end. This file is the reference.
Four stages, resumable at every step.
- transcribe — ffmpeg pulls a 16 kHz mono track, faster-whisper transcribes it locally. Game audio buries speech under music and effects, and a video model asked to transcribe it will invent lines. Whisper's output is passed to the annotator as ground truth so it quotes rather than guesses.
- split — a single encode pass cuts fixed-length segments, forcing keyframes at every boundary so cuts land on time.
- annotate — each segment goes to Gemini with its own slice of the transcript, rebased so the segment starts at 0, plus the prompt. The model returns structured JSON blocks.
- stitch — validates, adds each segment's offset back, renders the outputs.
Everything about a run lives in runs/<name>-<hash>/, keyed by the video's content,
so a job resumes even if you move or rename the file.
For how the code is laid out, see src/README.md.
python3 -m venv venv && source venv/bin/activate
pip install -e .
sudo apt install ffmpeg # ffmpeg and ffprobe are required
cp .env.example .env # then fill in credentialsTwo backends, resolved in this order:
- Vertex AI if forced (
VIDANNOTATE_USE_VERTEX=1orGOOGLE_GENAI_USE_VERTEXAI=true) and a project is configured. - Gemini Developer API if
GEMINI_API_KEYis set. - Vertex AI if a project is configured and there is no API key.
vidannotate backend # prints which one resolves, and whyVertex has no Files API, so segments are staged through Cloud Storage. Set
VIDANNOTATE_GCS_BUCKET and pip install -e '.[vertex]'. Without a bucket only
segments under about 18 MB can be sent inline, which at the default 90 s segment
length means effectively none.
vidannotate run video.mp4 # whole pipeline
vidannotate run video.mp4 --dry-run # plan only, nothing charged
vidannotate run video.mp4 --seg 60 --jobs 4 # shorter segments, more concurrency
vidannotate run video.mp4 --through split # stop after splitting
vidannotate run video.mp4 --only annotate # just that stage
vidannotate run video.mp4 --reset annotate # discard annotations and redo
vidannotate run video.mp4 --strict # refuse to stitch on validation issues
vidannotate status video.mp4 # progress, per stage and per segment
vidannotate redo video.mp4 -s 3 -s 5 # re-annotate specific segmentsInterrupt any run and re-issue the same command. Finished segments are skipped.
--prompt takes a bundled preset name or a path to a file, defaulting to
gameplay. A prompt says what to report; the output shape is fixed by the
response schema, so a prompt never needs formatting instructions. Writing one is
covered in src/vidannotate/prompts/README.md.
Before stitching, every segment is checked for:
- blocks out of order, overlapping, or reversed
- gaps in coverage
- a single block covering more than 30 seconds, meaning the model summarised a stretch rather than describing it beat by beat
- blocks falling outside the segment's own length, meaning the model was not working in the units it was asked for
- dialogue the model claims came from the transcript but that is not in it
Every quoted line carries a declared source of transcript or subtitle.
Transcript claims are verified against the transcript. Subtitle claims cannot be
checked from outside the video, but requiring the declaration keeps invented quotes
from arriving unmarked, and narrows spot-checking to a short list.
Results land in validation.json and print as a summary. --strict makes them
fatal.
The schema asks for seconds. Models tend to write video times as M:SS, and that
habit survives a numeric field: 1:30 arrives as 130. The value looks valid, so
the error is silent, and everything past the first minute of each segment ends up
misplaced by up to 40 seconds.
timecode.py detects the signature, which is values above the clip length combined
with an empty 60-99 band, and decodes them. It runs at annotation time and again at
stitch, is idempotent, and reports when it fires. This happened on the first
full-length run against real footage.
Whisper's VAD sometimes collapses a long silence into a single utterance and anchors
a short line across the whole span. Such a line would be injected into several
segments at wrong timestamps and quoted as dialogue that was never spoken there.
Anything longer than --max-utterance (default 15 s) is dropped and listed in
transcript.warnings.json. Fixing those lines by hand in transcript.txt and
re-running the affected segments recovers their coverage.
python -m unittest discover -s testsCovers transcript slicing and rebasing, artifact filtering, validation, and rendering: the pure logic where a bug corrupts timestamps without any visible sign.
MIT, see LICENSE.
This covers vidannotate itself, not its dependencies or the services it calls. faster-whisper and the Whisper models are MIT. Use of the Gemini API or Vertex AI is governed by Google's terms. Whatever video you run through it remains your responsibility.