Skip to content

Latest commit

 

History

11 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

vidannotate

Turns a video into a timestamped, action-by-action written description of what happens in it.

vidannotate run gameplay.mp4
# -> runs/gameplay-1a2b3c4d/final.txt   (also final.json, final.srt)

New here? GUIDE.md walks through a single video end to end. This file is the reference.

How it works

Four stages, resumable at every step.

  1. transcribe — ffmpeg pulls a 16 kHz mono track, faster-whisper transcribes it locally. Game audio buries speech under music and effects, and a video model asked to transcribe it will invent lines. Whisper's output is passed to the annotator as ground truth so it quotes rather than guesses.
  2. split — a single encode pass cuts fixed-length segments, forcing keyframes at every boundary so cuts land on time.
  3. annotate — each segment goes to Gemini with its own slice of the transcript, rebased so the segment starts at 0, plus the prompt. The model returns structured JSON blocks.
  4. stitch — validates, adds each segment's offset back, renders the outputs.

Everything about a run lives in runs/<name>-<hash>/, keyed by the video's content, so a job resumes even if you move or rename the file.

For how the code is laid out, see src/README.md.

Setup

python3 -m venv venv && source venv/bin/activate
pip install -e .
sudo apt install ffmpeg          # ffmpeg and ffprobe are required
cp .env.example .env             # then fill in credentials

Credentials

Two backends, resolved in this order:

  1. Vertex AI if forced (VIDANNOTATE_USE_VERTEX=1 or GOOGLE_GENAI_USE_VERTEXAI=true) and a project is configured.
  2. Gemini Developer API if GEMINI_API_KEY is set.
  3. Vertex AI if a project is configured and there is no API key.
vidannotate backend      # prints which one resolves, and why

Vertex has no Files API, so segments are staged through Cloud Storage. Set VIDANNOTATE_GCS_BUCKET and pip install -e '.[vertex]'. Without a bucket only segments under about 18 MB can be sent inline, which at the default 90 s segment length means effectively none.

Usage

vidannotate run video.mp4                    # whole pipeline
vidannotate run video.mp4 --dry-run          # plan only, nothing charged
vidannotate run video.mp4 --seg 60 --jobs 4  # shorter segments, more concurrency
vidannotate run video.mp4 --through split    # stop after splitting
vidannotate run video.mp4 --only annotate    # just that stage
vidannotate run video.mp4 --reset annotate   # discard annotations and redo
vidannotate run video.mp4 --strict           # refuse to stitch on validation issues

vidannotate status video.mp4                 # progress, per stage and per segment
vidannotate redo video.mp4 -s 3 -s 5         # re-annotate specific segments

Interrupt any run and re-issue the same command. Finished segments are skipped.

Prompts

--prompt takes a bundled preset name or a path to a file, defaulting to gameplay. A prompt says what to report; the output shape is fixed by the response schema, so a prompt never needs formatting instructions. Writing one is covered in src/vidannotate/prompts/README.md.

Validation

Before stitching, every segment is checked for:

  • blocks out of order, overlapping, or reversed
  • gaps in coverage
  • a single block covering more than 30 seconds, meaning the model summarised a stretch rather than describing it beat by beat
  • blocks falling outside the segment's own length, meaning the model was not working in the units it was asked for
  • dialogue the model claims came from the transcript but that is not in it

Every quoted line carries a declared source of transcript or subtitle. Transcript claims are verified against the transcript. Subtitle claims cannot be checked from outside the video, but requiring the declaration keeps invented quotes from arriving unmarked, and narrows spot-checking to a short list.

Results land in validation.json and print as a summary. --strict makes them fatal.

Timestamp units

The schema asks for seconds. Models tend to write video times as M:SS, and that habit survives a numeric field: 1:30 arrives as 130. The value looks valid, so the error is silent, and everything past the first minute of each segment ends up misplaced by up to 40 seconds.

timecode.py detects the signature, which is values above the clip length combined with an empty 60-99 band, and decodes them. It runs at annotation time and again at stitch, is idempotent, and reports when it fires. This happened on the first full-length run against real footage.

Known constraint

Whisper's VAD sometimes collapses a long silence into a single utterance and anchors a short line across the whole span. Such a line would be injected into several segments at wrong timestamps and quoted as dialogue that was never spoken there. Anything longer than --max-utterance (default 15 s) is dropped and listed in transcript.warnings.json. Fixing those lines by hand in transcript.txt and re-running the affected segments recovers their coverage.

Tests

python -m unittest discover -s tests

Covers transcript slicing and rebasing, artifact filtering, validation, and rendering: the pure logic where a bug corrupts timestamps without any visible sign.

License

MIT, see LICENSE.

This covers vidannotate itself, not its dependencies or the services it calls. faster-whisper and the Whisper models are MIT. Use of the Gemini API or Vertex AI is governed by Google's terms. Whatever video you run through it remains your responsibility.

About

An AI-powered CLI tool that annotates videos, and saves them in a transcript. Free to use, easy to set up.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages