audio.cpp ships one browser interface: a SvelteKit/TypeScript single-page app embedded directly in
audiocpp_server. The compiled server needs neither Python, Node.js, nor separate frontend files for
inference and normal UI operation.
Build audiocpp_server, then start it with UI management enabled:
.\build\windows-cuda-release\bin\audiocpp_server.exe --ui --backend cuda./build/bin/audiocpp_server --ui --backend cudaOpen http://127.0.0.1:8080. With no --config, --ui enables on-demand model loading,
unloading, package management, and temporary browser uploads. Models default to a models/ directory
beside the server executable. The Models page can select and remember a different directory.
An existing server configuration exposes the embedded UI by default:
audiocpp_server --config server.jsonConfigured models retain their eager or lazy behavior. Add --ui-management when the UI should be
allowed to load, switch, unload, download, or delete model packages. Use --no-ui for an API-only
server.
The native UI supports the shared model catalog and spec-driven controls for TTS, cloning, ASR, generation, conversion, separation, VAD, diarization, alignment, and voice design. It also provides:
- background model downloads with progress, cancellation, partial-download cleanup, version status, precision selection, and package deletion;
- on-demand model loading with automatic unloading when switching model or precision;
- sentence-aware long-text synthesis with browser-side WAV merging;
- microphone capture and near-live input for streaming-capable ASR models;
- embedded demo voices with matching transcripts;
- a saved voice library stored in browser IndexedDB;
- multilingual UI resources under
native/lang/; - structured results, generated artifacts, and request timing.
Uploaded request files use a per-process temporary directory and are deleted when the server exits. Saved voices remain in the current browser profile and are only uploaded when selected for a request.
Inference and the embedded interface do not require Python. Model installation currently invokes the
repository-level tools/model_manager_v2.py; legacy checkpoint conversion packages may invoke
tools/model_manager_deprecated.py. These are general model-management utilities, not UI code.
Set AUDIOCPP_PYTHON if the desired interpreter is not python on Windows or python3 on Unix.
Loading an existing model directory or standalone GGUF does not invoke either helper.
Node.js is needed only to modify and rebuild the frontend:
cd webui/native
npm ci
npm run check
npm run buildThe build creates webui/native/dist/index.html. CMake converts that single-file application into an
embedded byte array for audiocpp_server; rebuild the server after changing it. For live development,
run npm run dev; Vite proxies /health and /v1 to a server on port 8080.
The frontend consumes:
webui/configs/models_catalog.jsonfor model/task entries;webui/configs/model_params.jsonfor model-specific controls;webui/native/lang/lang_<code>.jsonfor optional translations;webui/native/demo_voices/for demo reference voices embedded in the server.
English strings are built into native/src/lib/i18n.ts and are the fallback for missing translations.