Skip to content

Latest commit

 

History

History

README.md

Data Processing Scripts

This directory contains the video, audio, and multimodal annotation scripts used to build the project dataset. All scripts read media paths from CSV files. Model weights and third-party source code are not included in this directory, and we recommend running scripts for different models in separate Python environments.

Script Overview

Script Description
video_scene_detection.py Detects video shot boundaries with PySceneDetect and outputs start/end timestamps and frame indices for video splitting.
annotate_video_motion.py Tracks points on a regular grid with CoTracker 3 and computes the average trajectory displacement between adjacent frames as a global motion metric.
score_video_aesthetics.py Samples video frames at fixed time intervals and computes an aesthetic score using CLIP and the MLP weights from improved-aesthetic-predictor.
detect_silent_audio.py Reads audio paths directly from a CSV file and determines whether each audio track is silent based on its overall loudness.
score_audio_quality.py Uses AudioBox Aesthetics to annotate four audio-quality dimensions: CE, CU, PC, and PQ.
classify_sound_events.py Uses Qwen3-Omni to identify the primary sound event and determine whether the media contains human speech.
detect_active_speakers.py Uses Fast-ASD/TalkNet to produce per-frame speaker bounding boxes and active-speaker scores.
transcribe_audio_elevenlabs.py Calls the commercial ElevenLabs Speech-to-Text API to transcribe audio in batches and saves the raw JSON responses.
segment_video_objects.py For videos without speech, detects objects with Grounding DINO and propagates them with SAM 2 to generate per-frame object masks.
segment_talking_people.py For videos with speech, combines Fast-ASD speaker boxes, Grounding DINO, and SAM 2 to generate per-frame masks for speaking people.
grounded_sam_utils.py Shared model-loading, video-propagation, and COCO RLE-encoding utilities for the two Grounded-SAM-2 annotation scripts.
mask_guided_editing_model/ A two-stage mask-guided editing pipeline that separates the target sound with SAM-Audio, generates a new sound effect, and produces an edited video with a Wan-based model. Custom weights are available at suimu/mask_guided_editing_model; the remaining Wan base components come from Wan2.2-TI2V-5B.

Dependencies and Upstream Projects

These scripts use or build upon the following open-source projects and models:

Before running the scripts, prepare the corresponding repositories, Python dependencies, and model weights, and comply with the licenses and model terms of use for each project. The ElevenLabs script reads credentials from a command-line argument or the ELEVENLABS_API_KEY environment variable; no API keys are stored in this repository.

The model-loading and inference workflow in score_audio_quality.py is adapted from Meta's AudioBox Aesthetics (CC BY 4.0). The other scripts use the public interfaces of their respective upstream projects. When distributing this directory, retain the project-level license and verify the licensing requirements for all upstream code and model weights.