Repository navigation
Tags: CyberSys/audio.cpp
Tags
moss_tts_v15: MOSS-TTS-v1.5, and promote the moss_tts_delay family ru… …ntimes (0xShug0#631) * moss codec: load a stage's input projection only when it exists Groundwork for the moss_tts_delay family (0xShug0#607), and a real fix on the way there. The v1 codec encoder has never run. moss_tts_local exercises v2 encode and moss_voicegen exercises v1 decode, but nothing calls v1 encode, so its encoder_stages had never executed. The delay-pattern models clone from a reference recording, which is what will call it. Running it fails at once: MOSS codec tensor not found: encoder.3.input_proj.weight. The checkpoint's projections are asymmetric -- encoder.1: input_proj, output_proj decoder.0: input_proj encoder.3: output_proj decoder.2: input_proj encoder.5: output_proj decoder.4: input_proj encoder.7: output_proj decoder.6: input_proj, output_proj -- because upstream's ProjectedTransformer only creates a projection when the stage changes width. The loader already knew this for output_proj, with a comment saying v1 omits it wherever output_dimension equals d_model, but loaded input_proj unconditionally. Against the v1 encoder stages the same width rule predicts the checkpoint exactly: stage 0: input 240 != d_model 768 -> encoder.1 has input_proj stage 1: input 768 == d_model 768 -> encoder.3 omits it stage 2: input 768 == d_model 768 -> encoder.5 omits it stage 3: input 1280 == d_model 1280 -> encoder.7 omits it so the config was right and only the loader was wrong. It survived because v1's decoder widens on input at every stage and narrows on output only at the last, so handling the output side alone satisfies the decoder completely. input_proj is now optional exactly as output_proj was, with the same width-change guard, and the graph passes the input through when it is absent. The change is in shared code, so v2 and nano inherit the rule. It only alters behaviour where the old code threw, so no previously passing path moves. Two parity tests, neither of which needs a GGUF conversion: the checkpoints' tensor names are already what the runtimes look up, so both read the HF download through open_tensor_source(). moss_voicegen_codec_encode_parity -- v1 encode, 3072/3072 codes exact moss_tts_delay_backbone_parity -- MOSS-TTS-v1.5, the 8B n_vq 32 sibling: prefill max_rel 0.0052, cached single step 0.0038, non_finite 0, passing at a tolerance of 0.006 The backbone test takes prompt rows from the reference dump instead of building them, because the v1.5 prompt builder does not exist yet; that isolates geometry, the summed multi-codebook embedding, rope and the KV cache from prompt construction, which is the split moss_voicegen already uses between prompt_parity and backbone_parity. A note for whoever runs these next: the RLFQ quantizer is sensitive enough that the reference disagrees with itself across devices. A CUDA reference and a CPU reference differ in 36 of 3072 codes at frames 63, 64, 75 and 78, which first read as a boundary bug in the port. Exact code match is a gate on the harness being device-matched as much as on the encoder. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01TMmMgd5xNGnjsQgybuUiK3 * moss: promote the delay-family runtimes into the framework The moss_tts_delay architecture covers several checkpoints that differ mainly in size and codebook count: MOSS-VoiceGenerator (Qwen3-1.7B, n_vq 16), MOSS-TTSD (8B, n_vq 16) and MOSS-TTS-v1.5 (8B, n_vq 32). Only moss_voicegen existed, so its backbone, heads and delay decoder were model-local although none of them is specific to that checkpoint. This follows 0xShug0#332, which promoted the MOSS audio tokenizer codec runtime into framework code so the family could share one lifecycle-managed runtime, and it is the only route available: no model in this tree includes another model's headers. moss_voicegen/{backbone,heads,delay_decoder}.{h,cpp} -> framework/decoders/moss_tts_delay/ (engine::decoders) new: framework/decoders/moss_tts_delay/config.h engine::decoders in framework/decoders is where tdt_decoder_* already lives, which is the existing example of a family-specific decode runtime in the framework. The geometry and token-id structs move out of moss_voicegen/assets.h into config.h; assets.h aliases them so that model's own call sites keep reading in its own vocabulary. One design change rather than a rename: the backbone and heads runtimes took a shared_ptr<const MossVoiceGenAssets> but used exactly two fields of it. They now take (MossTtsDelayConfig, shared_ptr<const TensorSource>). That is what lets a checkpoint with no MossVoiceGenAssets use them, which is the point of the promotion; the v1.5 backbone test had been constructing a fake assets struct purely to satisfy the old signature, and that scaffolding is gone. Profiling labels moved with the code, moss_voicegen.* -> moss_tts_delay.*. Nothing outside these files referenced them. Verified against MOSS-VoiceGenerator itself, not only against the checkpoint that motivated the move: prompt parity 4/4 fixtures, token for token backbone parity max_rel 3.20e-05 prefill, 1.78e-05 cached step generation parity 40/40 rows, 0 mismatching codec decode parity worst probe deviation 6.6e-07 smoke 6 s of audio, 75 frames, delay trace showing 74 gen slots then exactly n_vq=16 delay slots and the deviations on the 8B n_vq 32 path are unchanged to seven significant figures, 0.00520836 prefill and 0.00383948 cached step, which a refactor that altered behaviour would not reproduce. Full build clean; ctest 67/67. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01TMmMgd5xNGnjsQgybuUiK3 * moss_tts_v15: the prompt builder, with reference audio MOSS-TTS-v1.5's user turn is the moss_tts_delay family's eight-field <user_inst> block -- Reference(s), Instruction, Tokens, Quality, Sound Event, Ambient Sound, Language, Text. moss_voicegen already renders all eight but fills only the instruction and the language, leaving the rest at "None". v1.5 additionally uses the reference slot for cloning and the token budget for duration, so the fields are carried as options. A reference recording enters the prompt as its codec codes: the Reference(s) slot renders "[S<n>]:" and an audio span of audio_start, one audio_user_slot row per delayed frame, audio_end. The span length is frames + n_vq - 1, not frames + n_vq. Codebook v occupies row t + v, so the last row used is (frames - 1) + (n_vq - 1); the reference builds the same span as `length` gen slots followed by `n_vq - 1` delay slots. The first attempt was one row long and the parity test caught it as a row-count mismatch rather than as bad audio a stage later. Parity against dumps from the reference MossTTSDelayProcessor, comparing every channel rather than only the text one: instruction-only 83 rows, 0 text mismatches, 0 code mismatches voice clone 201 rows, 0 text mismatches, 0 code mismatches The cloning case is the one that matters: 201 x 32 code values, so the delay layout of the reference audio is checked position by position. The test compares the rendered <user_inst> before tokenising as well, because a template difference is readable as text and unreadable as a token-id mismatch. That caught a fixture bug on the first run -- the dump script wrote "language": "English" while the message it dumped had language=None, so the fixture described a prompt nobody had rendered. Fixtures now record the fields as the message actually carried them. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01TMmMgd5xNGnjsQgybuUiK3 * moss_tts_v15: the model, end to end MOSS-TTS-v1.5 now loads through the normal path and speaks. The family runtimes do the work; what is here is the checkpoint's own half -- assets, session, loader and spec -- plus one more promotion. The config parser moves to the family as well. It was 70 lines in moss_voicegen that read nothing but moss_tts_delay fields; only the error strings named a model, so it takes a label. moss_voicegen's assets.cpp drops from 87 lines to 33 and v1.5's is 20, where a copy would have been another 70. What is genuinely this checkpoint's: - the prompt, which fills the reference slot and the token budget that voice design leaves at "None" - encode_reference(), which downmixes and resamples a caller's recording to the codec's 24 kHz mono, trims to whole frames, and encodes it once per request rather than once per text chunk - a duration bound that prefers the model's own "- Tokens:" field and falls back to the character-rate estimate voice design has to use - tts and clone as one path, separated only by whether a reference recording is attached Verified through audiocpp_cli against the real checkpoint, transcribed to check the words rather than only the waveform: --task tts 3.52 s, "The quick brown fox jumps over the lazy dog." exact, median F0 117.3 Hz --task clon 3.68 s, exact, median F0 184.0 Hz against the reference recording's 189.4 -- and 117.3 without it, so the reference is doing the work Prompt parity holds for both shapes, 83 and 201 rows with zero text and zero code mismatches, and moss_voicegen's own parity is unchanged after the parser move. Full build clean; ctest 67/67. Two bugs found by running it, both mine. The session was copied from moss_voicegen and kept returning VoiceDesign from task_kind(), so a tts request reported task=vdes. And prepare() never constructed the tokenizer after the text processor was replaced by the free-standing builder, which segfaulted on the first encode -- a null this, not a subtle one, and it would have been caught by any run at all. Note for anyone adding a model spec: schema_version 1 validates far more than the older specs carry. It rejected 'voice_cloning' and 'duration_budget' as capabilities, required session options to use bare names with the family prefix derived, required weight_type to use its registered preset, required a 'dependencies' field, and rejected 'Voice Cloning' as a UI tag. The vocabularies are discoverable only by reading the other specs. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01TMmMgd5xNGnjsQgybuUiK3 * moss_tts_v15: GGUF packaging and docs The spec advertised a package that did not exist. It does now: the converter picks the family up from model_specs automatically, so no tooling was needed beyond the recipe. audiocpp_gguf \ --input model_weights=<root>/model.safetensors.index.json \ --input audio_tokenizer_weights=<root>/audio_tokenizer/model.safetensors.index.json \ --output moss_tts_v15_bf16_codec_f16.gguf \ --type bf16 --keep-type "audio_tokenizer_weights*=f16" \ --family moss_tts_v15 --root <root> 20.5 GB, 2063 tensors, both namespaces, the model spec and 22 sidecars embedded. Unlike moss_voicegen's package this carries the codec encoder as well as the decoder, because cloning needs it. Verified by synthesising from the GGUF rather than from the safetensors: cloning a 189.4 Hz reference returned 180.8 Hz with the transcript exact, against 184.0 Hz by the safetensors path. The two takes differ because bf16 and f16 conversion moves the logits enough to change which tokens are sampled; both track the reference and both say the words. One packaging note worth having in the docs rather than rediscovering: the converter requires audio_tokenizer/ to be a real directory. A symlink to the tokenizer snapshot resolves for every other tool here but makes the converter report the sidecar missing. The docs carry the limitation as a section rather than a footnote: voice-attribute instructions are followed only loosely, with the four measurements, the caveat that five samples of median F0 is a proxy rather than an evaluation, and a pointer to moss_voicegen for anyone who actually wants a voice from a written description. README and docs/tts.md list the model with the same caveat in one line. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01TMmMgd5xNGnjsQgybuUiK3 * moss_tts_v15: name the loader factory the way the catalog check expects check_loader_catalog_sync.py, which both CI workflows run before building, discovers loaders by matching make_<family>_loader. Adapting moss_voicegen's loader I had shortened its symbols -- make_loader, load_model, load_assets -- which reads more naturally inside a namespace that already says moss_tts_v15, but leaves the family invisible to the tooling: model_specs/moss_tts_v15.json has no registered loader family The check exits 1 on that, so CI would have failed on the spec I added two commits ago. Restored to the tree's convention: make_moss_tts_v15_loader, load_moss_tts_v15_model, load_moss_tts_v15_assets. check_loader_catalog_sync.py now exits 0, its --self-test passes, and the model still loads and clones after the rename. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01TMmMgd5xNGnjsQgybuUiK3 * text: port the MOSS-TTS input normaliser The reference normalises the text as it builds the user message (processing_moss_tts.py, build_user_message), so the prompt the model sees is the normalised one. Without this, markdown markers, repeated punctuation and CJK jammed against Latin all reach the model literally. Ported to codepoints and hand-rolled scanning rather than std::regex. The Python leans on lookbehind, which std::regex does not support at all, and on Unicode classes that byte-oriented matching gets wrong; the existing normalisers here work around the first by substituting \b, which does not survive the CJK rules. It is longer than the Python and it is the same algorithm stage for stage. The 38 vectors that ship inside the reference script are exported verbatim and kept as the contract, so it is upstream's definition of correct rather than one invented here. All 38 pass, and each expected output is also checked to be a fixed point -- the processor may normalise text that has already been normalised, and two of my first attempts were not idempotent. # Release notes - Fixed the .map leak - Shipped v2.3.1 See [the notes](https://example.com/release) for details!!! -> Release notes。Fixed the .map leak。Shipped v2.3.1。See the notes https://example.com/release for details! Three vectors failed on one cause worth recording: a bare number is not "latinish". "npm 包" keeps its space and "2026 年" closes up, so the rule that re-inserts a space on the closing side of a CJK boundary has to know whether the token it just emitted carried a Latin letter. Tracking that fixed the date, the clock time and the idempotence failures together. Clean prose is unchanged, which is why this could come last: the five prompts used to bring the model up pass through untouched, Chinese included, and prompt parity is unaffected. Full build clean; ctest 68/68; loader/catalog sync passes. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01TMmMgd5xNGnjsQgybuUiK3 * moss: fix eight issues found reviewing the branch **A hang, reachable from any caller's text.** restore_spans() replaced each ___PROTn___ placeholder by re-scanning the whole string until the token was gone. A protected span can contain placeholder text itself -- "https://x.com/___PROT0___" is a URL, so it is protected whole -- and restoring it puts the token back for the next iteration to find. The string grows without bound and the thread never returns. render_user_inst() normalises every request's text, so this wedged the session. Now one left-to-right pass that never re-scans what it has written. Confirmed before and after: the input hung under `timeout` and now returns. **A lone ellipsis became a full stop.** The reference alternation is (?:\.{3,}|…{2,}|……+); the port had `ellipses >= 1`, so "Hmm… ok." became "Hmm。ok." -- a trailing-off pause silently rewritten as a sentence break in the prompt the model sees. The shipped vectors all use two or more, so they missed it. Nine regression vectors added for both, with the expectations taken from the reference implementation rather than written by hand, and kept in their own file so upstream's 38 stay as they ship. **The repetition penalty ignored the prompt's audio.** The comment said the prompt rows are all pad so skipping them changes nothing. That was true of MOSS-VoiceGenerator, which has no audio in its prompt, and false here: a cloning prompt carries the reference recording's codes, so the codes most likely to recur were the only ones never penalised. The decoder now takes seed_prompt_codes(), kept apart from the generated history so extract_audio_codes() still returns only what was generated. moss_voicegen's generation parity is unchanged at 40/40, which is what distinguishes fixing the cloning path from breaking the shared decoder. **prepare() could report success with null members.** The guard returned early on backbone_ != nullptr, but heads_, codec_ and codebooks_ are built after it; a throw partway through left a retry taking the early return and run() dereferencing a null codec_. Guarded on the last member constructed. The comment above it was moss_voicegen's -- "no reference audio to encode", "5.7 GB model" -- and is now this model's. **--tokens was applied per chunk rather than per request.** A budget of 40 frames on a text that split three ways meant three takes of 40 frames each, roughly three times what was asked, each floored at 0.45 x 40 however short its own text. Now shared across chunks in proportion to their length. **Three smaller ones.** The spec declared the clone task but gave it no capabilities, where every other clone-capable spec here lists speaker_reference under it. dump_fixtures.py still wrote to tests/moss_tts_delay/reference, where nothing is committed any more, so re-running it refreshed nothing. And docs/community_models/models.md had no row for the model. Full build clean; ctest 69/69; catalog sync passes; prompt parity on both shapes, moss_voicegen prompt and generation parity all unchanged; cloning verified end to end with an exact transcript. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01TMmMgd5xNGnjsQgybuUiK3 * moss_tts_v15: keep a package's own weight type, and size the codec arenas Testing on CUDA for the first time turned up two things, one of which made the published quantisations pointless. **weight_storage_type_ defaulted to BF16, so a quantised package was dequantised on load.** q4_k and q8_0 therefore cost exactly as much memory as bf16 and the quantisation bought nothing. On a 3090, q4_k plain TTS peaks at 21.2 GiB forced to BF16 and 10.6 GiB left native. The default is Native now: a GGUF carries its own type, and safetensors are bf16 upstream so nothing changes there. The spec's session option default follows, with the reason written down rather than left as a value to be guessed at. That was also why cloning would not run on CUDA at any quantisation -- not the codec encoder being large, which is where the failure message pointed, but the baseline being twice what it should have been. Both quantisations clone on CUDA now: q4_k RTF 0.65 peak 14.1 GiB q8_0 RTF 0.77 peak 17.9 GiB **bf16 still cannot clone on a 24 GiB card**, and that one is real rather than a bug: cloning holds both halves of the codec resident where plain TTS needs only the decoder. Documented rather than worked around. **The codec graph arenas were moss_voicegen's 2 GiB, charged twice.** The codec runtime takes one arena for the encoder and one for the decoder, and this model holds both where moss_voicegen holds only the decoder. The graphs are a few hundred codec frames; 512 MiB is ample. Also worth recording: every RTF in this branch's history until now was wall-clock including model load, which for a 9 GB package is tens of seconds of disk I/O against about two seconds of work. session.wall_ms is the honest measure. CPU is ~16 RTF and CUDA ~0.65, so the port is faster than real time on a GPU and about 25x off it on this CPU -- a very different picture from the numbers I had been quoting. ctest 69/69; cloning verified on CUDA for both quantisations with exact transcripts and the voice still tracking the reference (q8_0 191.6 Hz, q4_k 174.7 Hz, reference 189.4 Hz). Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01TMmMgd5xNGnjsQgybuUiK3 * moss_tts_v15: run the text-normalisation tests in CI, and widen the RTF figures The normalisation tests were under ENGINE_BUILD_MODEL_TESTS, but CI sets only ENGINE_BUILD_TESTS, so the one test on this branch that CI could actually run was the one it never did. It needs no weights -- text in, text out, 47 vectors, milliseconds -- so it belongs with the tests that run everywhere. The parity harnesses stay under model tests because they genuinely need a checkpoint. Verified in a CI-shaped configuration (ENGINE_BUILD_TESTS only, no model tests) and through ci/run-local.sh linux-cpu, which now exists upstream: 38 tests, all passing. The RTF figures quoted a single run each. Three runs give 0.80-0.84 for q8_0 and 0.69-0.72 for q4_k, and a fourth q8_0 run had come back at 1.08, so a single decimal implied precision that is not there. Both the repo docs and the model card now give the range and say it is a band measured on an otherwise-in-use desktop GPU.