Tags: scriptease/audio.cpp
Tags
feat(model): add Confucius4-R2T2 real-time streaming ASR (r2t2_asr) (0… …xShug0#604) * feat(model): add Confucius4-R2T2 real-time streaming ASR (r2t2_asr) Port NetEase Youdao Confucius4-R2T2 (a Qwen3-ASR-1.7B fine-tune with Longest Stable Prefix decoding) as a community model family. Runtime * src/community_models/r2t2_asr + include/engine/community_models/r2t2_asr: assets, Whisper log-mel frontend, windowed audio tower, LSP streaming state machine, and the R2T2 text pipeline (punctuation-by-context, repetition repair, language-tag parsing, "|" truncation). * Schema-v1 spec-backed loader: model_specs/r2t2_asr.json is the single source of truth for metadata, capabilities, options and packages, and the loader factory lives next to the session, so the family ships no loader.{h,cpp}. * The thinker is a thin adapter over the shared runtime::GreedyQwenDecoderRuntime (graph lifetimes, audio-embedding injection, static KV cache); long audio uses engine::audio::plan_audio_chunks and the tower reuses the shared modules plus core::ensure_backend_addressable_layout. No framework files are modified. * Offline and append-only streaming modes. Committed deltas follow the upstream WebSocket integrator contract (slice by code-point length, authoritative final transcript); chunk sizes 80-2000 ms via r2t2_asr.chunk_size_ms. * Language accepts ISO-639 codes (zh/en/ja/...) or canonical prompt names (Chinese/English/...); Auto detects. Checkpoints below Q8_0 are rejected at load time with an actionable error. Packaging and UI * model_specs/r2t2_asr.json (schema v1, status community) with the HF safetensors package plus published Q8_0 and F16 GGUF packages (davidxifeng/Confucius4-R2T2-gguf). Each GGUF embeds its sidecars and the schema-v1 option contract, so one file is a fully standalone package. * webui: catalog entry, r2t2_asr session controls, GGUF install choices, live microphone transcription, and the language dropdown. * docs/community_models/r2t2.md, the ASR model table, app/server/example.json. Verification (macOS MPS reference from the upstream repository) * tests/r2t2_asr: MPS golden generator, per-chunk comparison harness, and a repo-native offline+streaming smoke test. * Chinese reference clip: offline text, committed delta stream and final transcript are exact, 21/21 committed prefixes identical (auto and forced language). * English clip: offline, committed stream and final transcript exact; two internal prefixes differ by one token, which the reference itself reproduces between its own fp16 and bf16 runs. * Q8_0 and F16 GGUF reproduce the same transcripts; 48 kHz stereo input goes through the shared mono conversion/resampling helper with identical results; server offline, SSE streaming and live-PCM paths verified; model unload returns RSS from 6.77 GiB to 0.20 GiB. * fix(r2t2): exclude rollback metadata from streaming deltas * refactor(r2t2): rename family to confucius4_r2t2 * docs(r2t2): track streaming graph reuse follow-up * docs(r2t2): keep release-specific GGUF details on the model card * perf(r2t2): reuse streaming encoder, prefill and decode graphs Streaming previously rebuilt every major graph per chunk: the encoder matched the exact accumulated frame count, the thinker prefill matched the prompt length, and the decode graph was replaced whenever the growing prompt outgrew its KV capacity. - Audio encoder gains an opt-in capacity-bucketed graph (1-2 chunks exact, then 4-chunk steps) with a per-run attention-mask refill for the valid token prefix; the offline path keeps exact per-frame graphs. - GreedyQwenDecoderRuntime::generate routes reuse_graphs requests through one reserved block-prefill graph (64-token blocks) and one decode graph with KV capacity grown in 128-token buckets, plus a fixed-width prompt embedding lookup graph; new prompts clear KV on device. - Add block-prefill graph build timing traces. On the M3 Metal 44-chunk golden clip this cuts graph builds from 44/44/44 to 8/13/3, graph-build time from ~1.06 s to ~0.10 s, stream wall time from 25.25 s to 23.67 s, and peak footprint from 7.32 GB to 7.14 GB with identical transcripts. test_confucius4_r2t2_graph_reuse pins the semantics: bit-exact encoder output when a bucket adds no padded tokens, a 2e-2 relative-RMSE noise budget plus run-to-run determinism when it does (ggml reduction order changes with sequence length), and decoder token parity across bucket growth and shrink. * fix(r2t2): guard Metal attention precision when reusing encoder graphs * docs(r2t2): clarify graph reuse validation and numerical limits * fix(r2t2): make the graph-reuse drift alarm backend-specific Tightening the padded-encoder alarm to 2e-3 alongside the Metal-only 64-token precision guard broke the test on CPU, its default backend: CPU has no kernel boundary to guard, and its fp32 reduction-order noise reaches about 1.9e-2 relative RMSE at 63 padded tokens, so the suite aborted at frames=101 before reaching the joint token checks. Split the alarm per backend: 2e-3 on Metal, where the 64-token capacity guard holds drift at or below 7.6e-4, and 2.5e-2 on CPU, above the measured 1.9e-2 ceiling. Verified on this machine: both backends pass the full suite, including joint encoder/decoder token parity on real audio prefixes with automatic and forced-English prompts. Also correct the r2t2 doc note that described the alarm as a single 2e-3 bound.
Tighten BreezeTTS streaming path (0xShug0#483) * Tighten BreezeTTS streaming path * Make BreezeTTS streaming incremental by default * Document BreezeTTS incremental streaming controls * Clean up Breeze streaming errors
Fix Chatterbox request sequence voice references (0xShug0#379)
ci: add tag-driven prebuilt release pipeline (0xShug0#286) * ci: add GitHub Actions release workflow [no release] * ci: fix release jobs (explicit win targets, install libomp on mac) [no release] * ci: only publish release when all backend jobs succeed [no release] * ci: provision mac OpenMP off, Windows Vulkan SDK + CUDA toolkit [no release] * ci: adopt llama.cpp release patterns (7z packing, CUDA 12/13 matrix, Vulkan SDK fix, OpenMP-off CUDA build) [no release] * ci: use valid Vulkan SDK 1.4.357.0 and full CUDA versions in matrix [no release] * ci: bump cuda-toolkit to v0.2.36, CUDA 12.4/13.2 [no release] * ci: source vcvarsall before CUDA cmake so nvcc finds cl.exe [no release] * ci: fix robocopy exit-code mapping in CUDA bundle step [no release] * ci: fix CUDA artifact zip path (build/bin is 1 level shallower than CPU/Vulkan preset bins) [no release] * ci: bundle CUDA runtime from bin/lib/bin-x64 and stop double-shipping runtime DLLs in CUDA release * ci(win-cuda): build with GGML_BACKEND_DL like llama.cpp so CUDA ships as ggml-cuda.dll instead of 3 monoliths * ci: port llama.cpp release pipeline for CPU/Vulkan/CUDA/Metal (get-tag-name, GGML_BACKEND_DL CUDA, robust cudart bundling) * ci(win-cuda): pin CMAKE_CUDA_ARCHITECTURES per toolkit; use CUDA 13.3 in the matrix * ci(win-cuda): use ggml-style virtual/real CUDA archs (faster) and bundle cufft64 runtime * ci: one-command releases (publish toggle, scripts/release.sh, docs/RELEASING.md) * ci: one-click GUI releases (publish defaults on) + GUI-first releasing docs * ci: remove CLI release helper; GUI-only releasing docs * ci: semver tag-driven releases (remove b<N>/auto-push); manual version input + publish gate * ci: preserve pre-release flag and notes when attaching binaries to a release * ci: upload only real package files (.zip/.tar.gz) to the release * ci: fix boolean publish gate (inputs.publish) so manual dispatch releases publish * ci: enable native model manager in prebuilt releases Add -DAUDIOCPP_BUILD_NATIVE_MODEL_MANAGER=ON to the CMAKE_ARGS so the prebuilt binaries ship with the self-contained native UI (model downloads, dynamic model management, etc.) instead of relying on external Python. Addresses: 0xShug0#286 (comment) * fix: pass native model manager flags to Windows builds too The env.CMAKE_ARGS is consumed only by the Linux and macOS jobs which call cmake directly. Windows CPU and Vulkan jobs use build_windows.ps1 which has its own CLI parameters (-DeploymentBuild, -NativeModelManager) and ignores CMAKE_ARGS. Windows CUDA jobs also hardcode cmake flags. This fix: - Adds -DeploymentBuild -NativeModelManager to all build_windows.ps1 calls - Adds -DAUDIOCPP_BUILD_NATIVE_MODEL_MANAGER=ON to the CUDA cmake command * fix: add native model manager support to all build scripts Port the upstream fix (0xShug0/audio.cpp@e9e8f14) to the fork: - Add --native-model-manager, --system-openssl, --boringssl-archive CLI flags - Add -DAUDIOCPP_BUILD_NATIVE_MODEL_MANAGER to cmake invocations - Print native model manager status in build output This completes the owner's request: the release CI now builds the self-contained native UI on all platforms. * fix: bundle tools/ and model_specs/ into prebuilt archives The server's model installer invokes tools/model_manager_v2.py for package downloads. Without it in the prebuilt zip, model download from the WebUI fails with 'model preparation helper was not found'. Include tools/ and model_specs/ at the archive root so the server can resolve its repository_root (traverses upward from the exe). * ci: upload raw bundle artifacts, split bin/cudart, archive only at release - Build jobs now upload the unpacked bundle (binaries + tools/ + model_specs/) instead of a zip-wrapped-in-an-artifact, so downloading an Actions artifact yields a ready-to-run folder. - CUDA jobs upload bin and cudart as two separate artifacts (the old wildcard path merged both into one ~1 GB blob). - The release job downloads per-artifact folders, creates the final zip/tar.gz archives there, and uploads those to the Release -- asset layout stays identical to v0.6.0. * ci: bundle MSVC runtime DLLs into Windows prebuilt packages Windows CPU/Vulkan/CUDA prebuilts link the dynamic MSVC runtime (default /MD) plus vcomp140.dll (OpenMP, from -DENGINE_ENABLE_OPENMP=ON). These are not guaranteed on clean/enterprise/Server/container Windows, so the packages bundle them app-locally (vcruntime140*.dll, msvcp140*.dll, and vcomp140.dll for CPU/Vulkan) from the toolchain's Redist\MSVC layout. Total added size is ~1 MB compressed - negligible vs the CPU/Vulkan/CUDA package sizes, and keeps installs self-contained (no VC++ Redistributable dependency). CUDA skips vcomp140 because the CUDA build disables OpenMP.
PreviousNext