This is the repo for "DiffAnon: Diffusion-based Prosody Control for Voice Anonymization".
This repo contains the DiffAnon training and inference pipeline using classifier-free guidance (CFG) for controlling source prosody preservation during voice anonymization.
You can check the Demo Webpage for speech samples.
Requirements
- Python 3.9+
- PyTorch + torchaudio
- See
requirements.txtfor the full environment used originally.
Data Preparation
This pipeline expects precomputed features alongside each .wav:
*.st_first.ptand*.st_all.ptfirst-level codec embeddings from SpeechTokenizer*.freevcs.ptspeaker embeddings (FreeVC speaker encoder)*.mpm.pt(prosody features)
You can generate these with:
python preprocess_anon.py --in_dir /path/to/wavs --out_dir /path/to/output --freevc_emb --speechtokenizer --mpmTraining
python train_anon.py --config_path config_anon.jsonOptional resume:
python train_anon.py --config_path config_anon.json --resume_dir /path/to/logs --resume_milestone 100Inference (directory, CFG)
python infer_anon_dir_cfg_randref.py \
--model_path /path/to/checkpoint.pt \
--config_path config_anon.json \
--input_dir /path/to/input_wavs \
--output_dir /path/to/output \
--mode anon_poolNotes
--modecontrols which conditioning signals are used (e.g.,null,resynt,anon,anon_pool,random).anon_poolexpects pooled embeddings (*.freevcs.pt) under the--anon_ref_rootdirectory.- CFG sampling is configured inside the inference scripts (
sampling='ddim_cfg',cfg_scales=[0.8, 0]).
Files
- Training entrypoint:
train_anon.py - Model:
model_anon.py - Dataset:
dataset_anon.py - Inference:
infer_anon_dir_cfg_randref.py,infer_anon_list_cfg.py
Credits
- NS2VC: https://github.com/adelacvg/NS2VC
- SpeechTokenizer: https://github.com/ZhangXInFD/SpeechTokenizer
- Masked Prosody Model (MPM): https://github.com/MiniXC/masked_prosody_model
- FreeVC: https://github.com/OlaWod/FreeVC