Skip to content

Tags: CyberSys/audio.cpp

Tags

last-docker-build

Toggle last-docker-build's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
moss_tts_v15: MOSS-TTS-v1.5, and promote the moss_tts_delay family ru…

…ntimes (0xShug0#631)

* moss codec: load a stage's input projection only when it exists

Groundwork for the moss_tts_delay family (0xShug0#607), and a real fix on the way
there.

The v1 codec encoder has never run. moss_tts_local exercises v2 encode and
moss_voicegen exercises v1 decode, but nothing calls v1 encode, so its
encoder_stages had never executed. The delay-pattern models clone from a
reference recording, which is what will call it.

Running it fails at once: MOSS codec tensor not found:
encoder.3.input_proj.weight. The checkpoint's projections are asymmetric --

    encoder.1: input_proj, output_proj     decoder.0: input_proj
    encoder.3: output_proj                 decoder.2: input_proj
    encoder.5: output_proj                 decoder.4: input_proj
    encoder.7: output_proj                 decoder.6: input_proj, output_proj

-- because upstream's ProjectedTransformer only creates a projection when the
stage changes width. The loader already knew this for output_proj, with a
comment saying v1 omits it wherever output_dimension equals d_model, but
loaded input_proj unconditionally. Against the v1 encoder stages the same
width rule predicts the checkpoint exactly:

    stage 0: input 240  != d_model 768   -> encoder.1 has input_proj
    stage 1: input 768  == d_model 768   -> encoder.3 omits it
    stage 2: input 768  == d_model 768   -> encoder.5 omits it
    stage 3: input 1280 == d_model 1280  -> encoder.7 omits it

so the config was right and only the loader was wrong. It survived because
v1's decoder widens on input at every stage and narrows on output only at the
last, so handling the output side alone satisfies the decoder completely.

input_proj is now optional exactly as output_proj was, with the same
width-change guard, and the graph passes the input through when it is absent.
The change is in shared code, so v2 and nano inherit the rule. It only alters
behaviour where the old code threw, so no previously passing path moves.

Two parity tests, neither of which needs a GGUF conversion: the checkpoints'
tensor names are already what the runtimes look up, so both read the HF
download through open_tensor_source().

  moss_voicegen_codec_encode_parity  -- v1 encode, 3072/3072 codes exact
  moss_tts_delay_backbone_parity     -- MOSS-TTS-v1.5, the 8B n_vq 32 sibling:
                                        prefill max_rel 0.0052, cached single
                                        step 0.0038, non_finite 0, passing at
                                        a tolerance of 0.006

The backbone test takes prompt rows from the reference dump instead of
building them, because the v1.5 prompt builder does not exist yet; that
isolates geometry, the summed multi-codebook embedding, rope and the KV cache
from prompt construction, which is the split moss_voicegen already uses
between prompt_parity and backbone_parity.

A note for whoever runs these next: the RLFQ quantizer is sensitive enough
that the reference disagrees with itself across devices. A CUDA reference and
a CPU reference differ in 36 of 3072 codes at frames 63, 64, 75 and 78, which
first read as a boundary bug in the port. Exact code match is a gate on the
harness being device-matched as much as on the encoder.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01TMmMgd5xNGnjsQgybuUiK3

* moss: promote the delay-family runtimes into the framework

The moss_tts_delay architecture covers several checkpoints that differ
mainly in size and codebook count: MOSS-VoiceGenerator (Qwen3-1.7B,
n_vq 16), MOSS-TTSD (8B, n_vq 16) and MOSS-TTS-v1.5 (8B, n_vq 32).
Only moss_voicegen existed, so its backbone, heads and delay decoder
were model-local although none of them is specific to that checkpoint.

This follows 0xShug0#332, which promoted the MOSS audio tokenizer codec runtime
into framework code so the family could share one lifecycle-managed
runtime, and it is the only route available: no model in this tree
includes another model's headers.

    moss_voicegen/{backbone,heads,delay_decoder}.{h,cpp}
      -> framework/decoders/moss_tts_delay/            (engine::decoders)
    new: framework/decoders/moss_tts_delay/config.h

engine::decoders in framework/decoders is where tdt_decoder_* already
lives, which is the existing example of a family-specific decode runtime
in the framework. The geometry and token-id structs move out of
moss_voicegen/assets.h into config.h; assets.h aliases them so that
model's own call sites keep reading in its own vocabulary.

One design change rather than a rename: the backbone and heads runtimes
took a shared_ptr<const MossVoiceGenAssets> but used exactly two fields
of it. They now take (MossTtsDelayConfig, shared_ptr<const
TensorSource>). That is what lets a checkpoint with no
MossVoiceGenAssets use them, which is the point of the promotion; the
v1.5 backbone test had been constructing a fake assets struct purely to
satisfy the old signature, and that scaffolding is gone.

Profiling labels moved with the code, moss_voicegen.* ->
moss_tts_delay.*. Nothing outside these files referenced them.

Verified against MOSS-VoiceGenerator itself, not only against the
checkpoint that motivated the move:

    prompt parity        4/4 fixtures, token for token
    backbone parity      max_rel 3.20e-05 prefill, 1.78e-05 cached step
    generation parity    40/40 rows, 0 mismatching
    codec decode parity  worst probe deviation 6.6e-07
    smoke                6 s of audio, 75 frames, delay trace showing
                         74 gen slots then exactly n_vq=16 delay slots

and the deviations on the 8B n_vq 32 path are unchanged to seven
significant figures, 0.00520836 prefill and 0.00383948 cached step,
which a refactor that altered behaviour would not reproduce.

Full build clean; ctest 67/67.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01TMmMgd5xNGnjsQgybuUiK3

* moss_tts_v15: the prompt builder, with reference audio

MOSS-TTS-v1.5's user turn is the moss_tts_delay family's eight-field
<user_inst> block -- Reference(s), Instruction, Tokens, Quality, Sound
Event, Ambient Sound, Language, Text. moss_voicegen already renders all
eight but fills only the instruction and the language, leaving the rest
at "None". v1.5 additionally uses the reference slot for cloning and the
token budget for duration, so the fields are carried as options.

A reference recording enters the prompt as its codec codes: the
Reference(s) slot renders "[S<n>]:" and an audio span of audio_start,
one audio_user_slot row per delayed frame, audio_end.

The span length is frames + n_vq - 1, not frames + n_vq. Codebook v
occupies row t + v, so the last row used is (frames - 1) + (n_vq - 1);
the reference builds the same span as `length` gen slots followed by
`n_vq - 1` delay slots. The first attempt was one row long and the
parity test caught it as a row-count mismatch rather than as bad audio
a stage later.

Parity against dumps from the reference MossTTSDelayProcessor, comparing
every channel rather than only the text one:

    instruction-only   83 rows,  0 text mismatches, 0 code mismatches
    voice clone       201 rows,  0 text mismatches, 0 code mismatches

The cloning case is the one that matters: 201 x 32 code values, so the
delay layout of the reference audio is checked position by position.

The test compares the rendered <user_inst> before tokenising as well,
because a template difference is readable as text and unreadable as a
token-id mismatch. That caught a fixture bug on the first run -- the
dump script wrote "language": "English" while the message it dumped had
language=None, so the fixture described a prompt nobody had rendered.
Fixtures now record the fields as the message actually carried them.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01TMmMgd5xNGnjsQgybuUiK3

* moss_tts_v15: the model, end to end

MOSS-TTS-v1.5 now loads through the normal path and speaks. The family
runtimes do the work; what is here is the checkpoint's own half --
assets, session, loader and spec -- plus one more promotion.

The config parser moves to the family as well. It was 70 lines in
moss_voicegen that read nothing but moss_tts_delay fields; only the error
strings named a model, so it takes a label. moss_voicegen's assets.cpp
drops from 87 lines to 33 and v1.5's is 20, where a copy would have been
another 70.

What is genuinely this checkpoint's:

  - the prompt, which fills the reference slot and the token budget that
    voice design leaves at "None"
  - encode_reference(), which downmixes and resamples a caller's
    recording to the codec's 24 kHz mono, trims to whole frames, and
    encodes it once per request rather than once per text chunk
  - a duration bound that prefers the model's own "- Tokens:" field and
    falls back to the character-rate estimate voice design has to use
  - tts and clone as one path, separated only by whether a reference
    recording is attached

Verified through audiocpp_cli against the real checkpoint, transcribed
to check the words rather than only the waveform:

    --task tts    3.52 s, "The quick brown fox jumps over the lazy dog."
                  exact, median F0 117.3 Hz
    --task clon   3.68 s, exact, median F0 184.0 Hz against the
                  reference recording's 189.4 -- and 117.3 without it,
                  so the reference is doing the work

Prompt parity holds for both shapes, 83 and 201 rows with zero text and
zero code mismatches, and moss_voicegen's own parity is unchanged after
the parser move. Full build clean; ctest 67/67.

Two bugs found by running it, both mine. The session was copied from
moss_voicegen and kept returning VoiceDesign from task_kind(), so a tts
request reported task=vdes. And prepare() never constructed the
tokenizer after the text processor was replaced by the free-standing
builder, which segfaulted on the first encode -- a null this, not a
subtle one, and it would have been caught by any run at all.

Note for anyone adding a model spec: schema_version 1 validates far more
than the older specs carry. It rejected 'voice_cloning' and
'duration_budget' as capabilities, required session options to use bare
names with the family prefix derived, required weight_type to use its
registered preset, required a 'dependencies' field, and rejected
'Voice Cloning' as a UI tag. The vocabularies are discoverable only by
reading the other specs.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01TMmMgd5xNGnjsQgybuUiK3

* moss_tts_v15: GGUF packaging and docs

The spec advertised a package that did not exist. It does now: the
converter picks the family up from model_specs automatically, so no
tooling was needed beyond the recipe.

    audiocpp_gguf \
      --input model_weights=<root>/model.safetensors.index.json \
      --input audio_tokenizer_weights=<root>/audio_tokenizer/model.safetensors.index.json \
      --output moss_tts_v15_bf16_codec_f16.gguf \
      --type bf16 --keep-type "audio_tokenizer_weights*=f16" \
      --family moss_tts_v15 --root <root>

20.5 GB, 2063 tensors, both namespaces, the model spec and 22 sidecars
embedded. Unlike moss_voicegen's package this carries the codec encoder
as well as the decoder, because cloning needs it.

Verified by synthesising from the GGUF rather than from the safetensors:
cloning a 189.4 Hz reference returned 180.8 Hz with the transcript
exact, against 184.0 Hz by the safetensors path. The two takes differ
because bf16 and f16 conversion moves the logits enough to change which
tokens are sampled; both track the reference and both say the words.

One packaging note worth having in the docs rather than rediscovering:
the converter requires audio_tokenizer/ to be a real directory. A
symlink to the tokenizer snapshot resolves for every other tool here but
makes the converter report the sidecar missing.

The docs carry the limitation as a section rather than a footnote:
voice-attribute instructions are followed only loosely, with the four
measurements, the caveat that five samples of median F0 is a proxy
rather than an evaluation, and a pointer to moss_voicegen for anyone who
actually wants a voice from a written description. README and docs/tts.md
list the model with the same caveat in one line.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01TMmMgd5xNGnjsQgybuUiK3

* moss_tts_v15: name the loader factory the way the catalog check expects

check_loader_catalog_sync.py, which both CI workflows run before
building, discovers loaders by matching make_<family>_loader. Adapting
moss_voicegen's loader I had shortened its symbols -- make_loader,
load_model, load_assets -- which reads more naturally inside a namespace
that already says moss_tts_v15, but leaves the family invisible to the
tooling:

    model_specs/moss_tts_v15.json has no registered loader family

The check exits 1 on that, so CI would have failed on the spec I added
two commits ago. Restored to the tree's convention:
make_moss_tts_v15_loader, load_moss_tts_v15_model,
load_moss_tts_v15_assets.

check_loader_catalog_sync.py now exits 0, its --self-test passes, and
the model still loads and clones after the rename.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01TMmMgd5xNGnjsQgybuUiK3

* text: port the MOSS-TTS input normaliser

The reference normalises the text as it builds the user message
(processing_moss_tts.py, build_user_message), so the prompt the model
sees is the normalised one. Without this, markdown markers, repeated
punctuation and CJK jammed against Latin all reach the model literally.

Ported to codepoints and hand-rolled scanning rather than std::regex.
The Python leans on lookbehind, which std::regex does not support at
all, and on Unicode classes that byte-oriented matching gets wrong; the
existing normalisers here work around the first by substituting \b,
which does not survive the CJK rules. It is longer than the Python and
it is the same algorithm stage for stage.

The 38 vectors that ship inside the reference script are exported
verbatim and kept as the contract, so it is upstream's definition of
correct rather than one invented here. All 38 pass, and each expected
output is also checked to be a fixed point -- the processor may
normalise text that has already been normalised, and two of my first
attempts were not idempotent.

    # Release notes
    - Fixed the .map leak
    - Shipped v2.3.1
    See [the notes](https://example.com/release) for details!!!
    ->
    Release notes。Fixed the .map leak。Shipped v2.3.1。See the notes
    https://example.com/release for details!

Three vectors failed on one cause worth recording: a bare number is not
"latinish". "npm 包" keeps its space and "2026 年" closes up, so the rule
that re-inserts a space on the closing side of a CJK boundary has to
know whether the token it just emitted carried a Latin letter. Tracking
that fixed the date, the clock time and the idempotence failures
together.

Clean prose is unchanged, which is why this could come last: the five
prompts used to bring the model up pass through untouched, Chinese
included, and prompt parity is unaffected.

Full build clean; ctest 68/68; loader/catalog sync passes.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01TMmMgd5xNGnjsQgybuUiK3

* moss: fix eight issues found reviewing the branch

**A hang, reachable from any caller's text.** restore_spans() replaced
each ___PROTn___ placeholder by re-scanning the whole string until the
token was gone. A protected span can contain placeholder text itself --
"https://x.com/___PROT0___" is a URL, so it is protected whole -- and
restoring it puts the token back for the next iteration to find. The
string grows without bound and the thread never returns. render_user_inst()
normalises every request's text, so this wedged the session. Now one
left-to-right pass that never re-scans what it has written. Confirmed
before and after: the input hung under `timeout` and now returns.

**A lone ellipsis became a full stop.** The reference alternation is
(?:\.{3,}|…{2,}|……+); the port had `ellipses >= 1`, so "Hmm… ok."
became "Hmm。ok." -- a trailing-off pause silently rewritten as a
sentence break in the prompt the model sees. The shipped vectors all use
two or more, so they missed it.

Nine regression vectors added for both, with the expectations taken from
the reference implementation rather than written by hand, and kept in
their own file so upstream's 38 stay as they ship.

**The repetition penalty ignored the prompt's audio.** The comment said
the prompt rows are all pad so skipping them changes nothing. That was
true of MOSS-VoiceGenerator, which has no audio in its prompt, and false
here: a cloning prompt carries the reference recording's codes, so the
codes most likely to recur were the only ones never penalised. The
decoder now takes seed_prompt_codes(), kept apart from the generated
history so extract_audio_codes() still returns only what was generated.
moss_voicegen's generation parity is unchanged at 40/40, which is what
distinguishes fixing the cloning path from breaking the shared decoder.

**prepare() could report success with null members.** The guard returned
early on backbone_ != nullptr, but heads_, codec_ and codebooks_ are
built after it; a throw partway through left a retry taking the early
return and run() dereferencing a null codec_. Guarded on the last member
constructed. The comment above it was moss_voicegen's -- "no reference
audio to encode", "5.7 GB model" -- and is now this model's.

**--tokens was applied per chunk rather than per request.** A budget of
40 frames on a text that split three ways meant three takes of 40 frames
each, roughly three times what was asked, each floored at 0.45 x 40
however short its own text. Now shared across chunks in proportion to
their length.

**Three smaller ones.** The spec declared the clone task but gave it no
capabilities, where every other clone-capable spec here lists
speaker_reference under it. dump_fixtures.py still wrote to
tests/moss_tts_delay/reference, where nothing is committed any more, so
re-running it refreshed nothing. And docs/community_models/models.md had
no row for the model.

Full build clean; ctest 69/69; catalog sync passes; prompt parity on both
shapes, moss_voicegen prompt and generation parity all unchanged; cloning
verified end to end with an exact transcript.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01TMmMgd5xNGnjsQgybuUiK3

* moss_tts_v15: keep a package's own weight type, and size the codec arenas

Testing on CUDA for the first time turned up two things, one of which
made the published quantisations pointless.

**weight_storage_type_ defaulted to BF16, so a quantised package was
dequantised on load.** q4_k and q8_0 therefore cost exactly as much
memory as bf16 and the quantisation bought nothing. On a 3090, q4_k
plain TTS peaks at 21.2 GiB forced to BF16 and 10.6 GiB left native.
The default is Native now: a GGUF carries its own type, and safetensors
are bf16 upstream so nothing changes there. The spec's session option
default follows, with the reason written down rather than left as a
value to be guessed at.

That was also why cloning would not run on CUDA at any quantisation --
not the codec encoder being large, which is where the failure message
pointed, but the baseline being twice what it should have been. Both
quantisations clone on CUDA now:

    q4_k   RTF 0.65   peak 14.1 GiB
    q8_0   RTF 0.77   peak 17.9 GiB

**bf16 still cannot clone on a 24 GiB card**, and that one is real
rather than a bug: cloning holds both halves of the codec resident where
plain TTS needs only the decoder. Documented rather than worked around.

**The codec graph arenas were moss_voicegen's 2 GiB, charged twice.**
The codec runtime takes one arena for the encoder and one for the
decoder, and this model holds both where moss_voicegen holds only the
decoder. The graphs are a few hundred codec frames; 512 MiB is ample.

Also worth recording: every RTF in this branch's history until now was
wall-clock including model load, which for a 9 GB package is tens of
seconds of disk I/O against about two seconds of work. session.wall_ms
is the honest measure. CPU is ~16 RTF and CUDA ~0.65, so the port is
faster than real time on a GPU and about 25x off it on this CPU -- a
very different picture from the numbers I had been quoting.

ctest 69/69; cloning verified on CUDA for both quantisations with exact
transcripts and the voice still tracking the reference (q8_0 191.6 Hz,
q4_k 174.7 Hz, reference 189.4 Hz).

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01TMmMgd5xNGnjsQgybuUiK3

* moss_tts_v15: run the text-normalisation tests in CI, and widen the RTF figures

The normalisation tests were under ENGINE_BUILD_MODEL_TESTS, but CI sets
only ENGINE_BUILD_TESTS, so the one test on this branch that CI could
actually run was the one it never did. It needs no weights -- text in,
text out, 47 vectors, milliseconds -- so it belongs with the tests that
run everywhere. The parity harnesses stay under model tests because they
genuinely need a checkpoint.

Verified in a CI-shaped configuration (ENGINE_BUILD_TESTS only, no model
tests) and through ci/run-local.sh linux-cpu, which now exists upstream:
38 tests, all passing.

The RTF figures quoted a single run each. Three runs give 0.80-0.84 for
q8_0 and 0.69-0.72 for q4_k, and a fourth q8_0 run had come back at 1.08,
so a single decimal implied precision that is not there. Both the repo
docs and the model card now give the range and say it is a band measured
on an otherwise-in-use desktop GPU.

release-0.4.2

Toggle release-0.4.2's commit message
Fix Qwen3 TTS decoder chunk graph reuse

release-0.4

Toggle release-0.4's commit message
Release 0.4

release-0.3-qwen3-tts

Toggle release-0.3-qwen3-tts's commit message
Add Qwen3 TTS issue 67 warmbench case

release-0.3

Toggle release-0.3's commit message
Update release notes and contribution guidance

v0.2.0-windows-prebuilt

Toggle v0.2.0-windows-prebuilt's commit message
audio.cpp Windows prebuilt binaries v0.2.0

release-0.2

Toggle release-0.2's commit message
Extend conv transpose fast path unit test

v0.1.0-windows-prebuilt

Toggle v0.1.0-windows-prebuilt's commit message
Windows prebuilt binaries