Skip to content

Tags: scriptease/audio.cpp

Tags

last-docker-build

Toggle last-docker-build's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
feat(model): add Confucius4-R2T2 real-time streaming ASR (r2t2_asr) (0…

…xShug0#604)

* feat(model): add Confucius4-R2T2 real-time streaming ASR (r2t2_asr)

Port NetEase Youdao Confucius4-R2T2 (a Qwen3-ASR-1.7B fine-tune with Longest
Stable Prefix decoding) as a community model family.

Runtime
* src/community_models/r2t2_asr + include/engine/community_models/r2t2_asr:
  assets, Whisper log-mel frontend, windowed audio tower, LSP streaming state
  machine, and the R2T2 text pipeline (punctuation-by-context, repetition
  repair, language-tag parsing, "|" truncation).
* Schema-v1 spec-backed loader: model_specs/r2t2_asr.json is the single source
  of truth for metadata, capabilities, options and packages, and the loader
  factory lives next to the session, so the family ships no loader.{h,cpp}.
* The thinker is a thin adapter over the shared
  runtime::GreedyQwenDecoderRuntime (graph lifetimes, audio-embedding
  injection, static KV cache); long audio uses engine::audio::plan_audio_chunks
  and the tower reuses the shared modules plus
  core::ensure_backend_addressable_layout. No framework files are modified.
* Offline and append-only streaming modes. Committed deltas follow the upstream
  WebSocket integrator contract (slice by code-point length, authoritative
  final transcript); chunk sizes 80-2000 ms via r2t2_asr.chunk_size_ms.
* Language accepts ISO-639 codes (zh/en/ja/...) or canonical prompt names
  (Chinese/English/...); Auto detects. Checkpoints below Q8_0 are rejected at
  load time with an actionable error.

Packaging and UI
* model_specs/r2t2_asr.json (schema v1, status community) with the HF
  safetensors package plus published Q8_0 and F16 GGUF packages
  (davidxifeng/Confucius4-R2T2-gguf). Each GGUF embeds its sidecars and the
  schema-v1 option contract, so one file is a fully standalone package.
* webui: catalog entry, r2t2_asr session controls, GGUF install choices, live
  microphone transcription, and the language dropdown.
* docs/community_models/r2t2.md, the ASR model table, app/server/example.json.

Verification (macOS MPS reference from the upstream repository)
* tests/r2t2_asr: MPS golden generator, per-chunk comparison harness, and a
  repo-native offline+streaming smoke test.
* Chinese reference clip: offline text, committed delta stream and final
  transcript are exact, 21/21 committed prefixes identical (auto and forced
  language).
* English clip: offline, committed stream and final transcript exact; two
  internal prefixes differ by one token, which the reference itself reproduces
  between its own fp16 and bf16 runs.
* Q8_0 and F16 GGUF reproduce the same transcripts; 48 kHz stereo input goes
  through the shared mono conversion/resampling helper with identical results;
  server offline, SSE streaming and live-PCM paths verified; model unload
  returns RSS from 6.77 GiB to 0.20 GiB.

* fix(r2t2): exclude rollback metadata from streaming deltas

* refactor(r2t2): rename family to confucius4_r2t2

* docs(r2t2): track streaming graph reuse follow-up

* docs(r2t2): keep release-specific GGUF details on the model card

* perf(r2t2): reuse streaming encoder, prefill and decode graphs

Streaming previously rebuilt every major graph per chunk: the encoder
matched the exact accumulated frame count, the thinker prefill matched
the prompt length, and the decode graph was replaced whenever the growing
prompt outgrew its KV capacity.

- Audio encoder gains an opt-in capacity-bucketed graph (1-2 chunks
  exact, then 4-chunk steps) with a per-run attention-mask refill for
  the valid token prefix; the offline path keeps exact per-frame graphs.
- GreedyQwenDecoderRuntime::generate routes reuse_graphs requests
  through one reserved block-prefill graph (64-token blocks) and one
  decode graph with KV capacity grown in 128-token buckets, plus a
  fixed-width prompt embedding lookup graph; new prompts clear KV on
  device.
- Add block-prefill graph build timing traces.

On the M3 Metal 44-chunk golden clip this cuts graph builds from
44/44/44 to 8/13/3, graph-build time from ~1.06 s to ~0.10 s, stream
wall time from 25.25 s to 23.67 s, and peak footprint from 7.32 GB to
7.14 GB with identical transcripts.

test_confucius4_r2t2_graph_reuse pins the semantics: bit-exact encoder
output when a bucket adds no padded tokens, a 2e-2 relative-RMSE noise
budget plus run-to-run determinism when it does (ggml reduction order
changes with sequence length), and decoder token parity across bucket
growth and shrink.

* fix(r2t2): guard Metal attention precision when reusing encoder graphs

* docs(r2t2): clarify graph reuse validation and numerical limits

* fix(r2t2): make the graph-reuse drift alarm backend-specific

Tightening the padded-encoder alarm to 2e-3 alongside the Metal-only
64-token precision guard broke the test on CPU, its default backend:
CPU has no kernel boundary to guard, and its fp32 reduction-order
noise reaches about 1.9e-2 relative RMSE at 63 padded tokens, so the
suite aborted at frames=101 before reaching the joint token checks.

Split the alarm per backend: 2e-3 on Metal, where the 64-token
capacity guard holds drift at or below 7.6e-4, and 2.5e-2 on CPU,
above the measured 1.9e-2 ceiling. Verified on this machine: both
backends pass the full suite, including joint encoder/decoder token
parity on real audio prefixes with automatic and forced-English
prompts.

Also correct the r2t2 doc note that described the alarm as a single
2e-3 bound.

v0.8.1

Toggle v0.8.1's commit message
Release v0.8.1

v0.8.0

Toggle v0.8.0's commit message
Release v0.8.0

v0.7.4

Toggle v0.7.4's commit message
Release v0.7.4

v0.7.3

Toggle v0.7.3's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
Tighten BreezeTTS streaming path (0xShug0#483)

* Tighten BreezeTTS streaming path

* Make BreezeTTS streaming incremental by default

* Document BreezeTTS incremental streaming controls

* Clean up Breeze streaming errors

v0.7.2

Toggle v0.7.2's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
Fix CosyVoice3 Metal flow layout (0xShug0#455)

v0.7.1

Toggle v0.7.1's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
Fix Chatterbox request sequence voice references (0xShug0#379)

v0.7.0

Toggle v0.7.0's commit message
Release 0.7

v0.6.2-release-test

Toggle v0.6.2-release-test's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
ci: add tag-driven prebuilt release pipeline (0xShug0#286)

* ci: add GitHub Actions release workflow [no release]

* ci: fix release jobs (explicit win targets, install libomp on mac) [no release]

* ci: only publish release when all backend jobs succeed [no release]

* ci: provision mac OpenMP off, Windows Vulkan SDK + CUDA toolkit [no release]

* ci: adopt llama.cpp release patterns (7z packing, CUDA 12/13 matrix, Vulkan SDK fix, OpenMP-off CUDA build) [no release]

* ci: use valid Vulkan SDK 1.4.357.0 and full CUDA versions in matrix [no release]

* ci: bump cuda-toolkit to v0.2.36, CUDA 12.4/13.2 [no release]

* ci: source vcvarsall before CUDA cmake so nvcc finds cl.exe [no release]

* ci: fix robocopy exit-code mapping in CUDA bundle step [no release]

* ci: fix CUDA artifact zip path (build/bin is 1 level shallower than CPU/Vulkan preset bins) [no release]

* ci: bundle CUDA runtime from bin/lib/bin-x64 and stop double-shipping runtime DLLs in CUDA release

* ci(win-cuda): build with GGML_BACKEND_DL like llama.cpp so CUDA ships as ggml-cuda.dll instead of 3 monoliths

* ci: port llama.cpp release pipeline for CPU/Vulkan/CUDA/Metal (get-tag-name, GGML_BACKEND_DL CUDA, robust cudart bundling)

* ci(win-cuda): pin CMAKE_CUDA_ARCHITECTURES per toolkit; use CUDA 13.3 in the matrix

* ci(win-cuda): use ggml-style virtual/real CUDA archs (faster) and bundle cufft64 runtime

* ci: one-command releases (publish toggle, scripts/release.sh, docs/RELEASING.md)

* ci: one-click GUI releases (publish defaults on) + GUI-first releasing docs

* ci: remove CLI release helper; GUI-only releasing docs

* ci: semver tag-driven releases (remove b<N>/auto-push); manual version input + publish gate

* ci: preserve pre-release flag and notes when attaching binaries to a release

* ci: upload only real package files (.zip/.tar.gz) to the release

* ci: fix boolean publish gate (inputs.publish) so manual dispatch releases publish

* ci: enable native model manager in prebuilt releases

Add -DAUDIOCPP_BUILD_NATIVE_MODEL_MANAGER=ON to the CMAKE_ARGS so the
prebuilt binaries ship with the self-contained native UI (model downloads,
dynamic model management, etc.) instead of relying on external Python.

Addresses: 0xShug0#286 (comment)

* fix: pass native model manager flags to Windows builds too

The env.CMAKE_ARGS is consumed only by the Linux and macOS jobs which
call cmake directly. Windows CPU and Vulkan jobs use build_windows.ps1
which has its own CLI parameters (-DeploymentBuild, -NativeModelManager)
and ignores CMAKE_ARGS. Windows CUDA jobs also hardcode cmake flags.

This fix:
- Adds -DeploymentBuild -NativeModelManager to all build_windows.ps1 calls
- Adds -DAUDIOCPP_BUILD_NATIVE_MODEL_MANAGER=ON to the CUDA cmake command

* fix: add native model manager support to all build scripts

Port the upstream fix (0xShug0/audio.cpp@e9e8f14) to the fork:
- Add --native-model-manager, --system-openssl, --boringssl-archive CLI flags
- Add -DAUDIOCPP_BUILD_NATIVE_MODEL_MANAGER to cmake invocations
- Print native model manager status in build output

This completes the owner's request: the release CI now builds the
self-contained native UI on all platforms.

* fix: bundle tools/ and model_specs/ into prebuilt archives

The server's model installer invokes tools/model_manager_v2.py for
package downloads. Without it in the prebuilt zip, model download from
the WebUI fails with 'model preparation helper was not found'.

Include tools/ and model_specs/ at the archive root so the server
can resolve its repository_root (traverses upward from the exe).

* ci: upload raw bundle artifacts, split bin/cudart, archive only at release

- Build jobs now upload the unpacked bundle (binaries + tools/ +
  model_specs/) instead of a zip-wrapped-in-an-artifact, so downloading
  an Actions artifact yields a ready-to-run folder.
- CUDA jobs upload bin and cudart as two separate artifacts (the old
  wildcard path merged both into one ~1 GB blob).
- The release job downloads per-artifact folders, creates the final
  zip/tar.gz archives there, and uploads those to the Release -- asset
  layout stays identical to v0.6.0.

* ci: bundle MSVC runtime DLLs into Windows prebuilt packages

Windows CPU/Vulkan/CUDA prebuilts link the dynamic MSVC runtime (default
/MD) plus vcomp140.dll (OpenMP, from -DENGINE_ENABLE_OPENMP=ON). These
are not guaranteed on clean/enterprise/Server/container Windows, so the
packages bundle them app-locally (vcruntime140*.dll, msvcp140*.dll, and
vcomp140.dll for CPU/Vulkan) from the toolchain's Redist\MSVC layout.

Total added size is ~1 MB compressed - negligible vs the CPU/Vulkan/CUDA
package sizes, and keeps installs self-contained (no VC++ Redistributable
dependency). CUDA skips vcomp140 because the CUDA build disables OpenMP.

release-0.6.1-brew-test

Toggle release-0.6.1-brew-test's commit message
Reduce Supertonic vector graph arena