Repository navigation
candle-fork plan: ONNX/ViT/Whisper via ndarray + Pi/NEON edge matrix #136
Description
Activity
R-BHQA review: red
Eleven AC clusters in one issue. This is a multi-month workstream filed as a single ticket.
-
Decompose. Minimum split: (a) fork + manifest swap (
accelerate-src/cblas→AdaWorldAPI/ndarray); (b)candle-onnxrewire and ONNX clinical-BERT smoke; (c) ViT smoke; (d) Whisper smoke; (e) embedanything Jina5 smoke; (f) edge deployment matrix; (g) license whitelist policy; (h) upstream merge cadence. The fork+swap is mechanical; everything after is investigative work that should not be gated on the same PR. -
Edge matrix is its own multi-week investigation. "Pi 5 (NEON, 8GB), Pi Zero 2 W (NEON, 512MB), Orange Pi (NEON, varies), x86 SPR (AMX) — each target builds and runs ViT inference." A 512MB Pi Zero 2 W running ViT-Base is not a smoke check, it's a feasibility study. Pull it out.
-
D1's tolerances do not carry over. AC for the four smoke tests says thresholds are "the tolerance defined by D1's parity harness." D1 (Numerical parity harness: bit-exact f32/f16/bf16 vs numpy/MKL/OpenBLAS/upstream #137) defines per-kernel ULP bands for gemm/add/mul/sum/conv2d. It does not define top-K embedding agreement, BLEU/CER for Whisper, or retrieval@K for Jina5. Those are model-level metrics that require their own thresholds. Either inline the metric thresholds here, or file an explicit dep on a separate "model-level tolerance" issue.
-
License audit is a compliance workstream. "Whitelist policy: model weights, license SPDX, commercial-use status, in-scope-for-clinical decision, last-audited date" is multi-stakeholder governance work, not an AC bullet on a fork-creation ticket.
-
Repo creation is not "execution under this plan." Creating
AdaWorldAPI/candlehas org-permissions implications; folding it into the issue's execution scope obscures the governance step. Make it an explicit prerequisite. -
Merge cadence is ongoing operational work. "Weekly
git fetch upstream+ integration-branch rebase, conflict resolution log, regression-gate" — this is recurring work tracked under a one-shot ticket. File as a separate process item in the candle-fork repo once it exists. -
NEON-Pi-5 ViT thresholds. Same model on SPR-AMX vs Pi-5-NEON must agree to within a defended tolerance — that is the entire moat statement in the issue body. Specify the tolerance here, not "TBD."
Generated by Claude Code
-
R-SHSR — most valuable single refinement: split this issue into MVP (fork + manifest swap + ONNX MiniLM smoke on x86_64-linux) plus 6 follow-up issues for the rest. Current scope is a quarter of work.
Suggested MVP (~1 week, single PR):
- Create
AdaWorldAPI/candlefork; tag upstream commit; integration branch as default. - Manifest swap:
accelerate-srcandcblas→AdaWorldAPI/ndarray. Remove the platformcfggates. - One smoke:
sentence-transformers/all-MiniLM-L6-v2ONNX (22 MB, MIT, deterministic, well-known reference embeddings). Top-K embedding agreement vsonnxruntimeon 5 sentences. - Build verification on
x86_64-linuxonly.
Defer to follow-up issues:
- F1: ViT smoke (ViT-Base parity)
- F2: Whisper smoke (German clinical clip)
- F3: embedanything Jina5 retrieval
- F4: Pi 5 NEON edge target
- F5: Pi Zero 2 W + Orange Pi targets
- F6: License whitelist + merge-cadence doc + German clinical-BERT licence audit
Why MiniLM not German-clinical-BERT for the MVP: sidesteps the clinical licence audit on Day 1; deterministic; widely-known reference outputs; <25 MB so weights ship in CI.
Strategic note: the "one binary, two accelerators" moat (SPR-AMX cloud + Pi 5 NEON edge) is real and worth preserving — but it is proven on the second PR (F4 Pi 5 NEON), not the MVP. Don't gate the MVP on edge validation; gate F4 on it.
Generated by Claude Code
- Create
Cross-reference review (relocated from issue body)
This review was originally appended to the issue body by an orchestration agent that lacked the
add_issue_commenttool. Relocating to a proper comment so the body reflects the original spec.
M-S5 cross-reference (model frameworks surface, 2026-05-04)
E2 blocks on E1 ndarray#135 (explicit in Dependencies — good). Don't fork candle before burn proves the spine pattern.
Pi 5 / Pi Zero 2 W edge deployment matrix realistically scopes to Q4_K_M-or-smaller for LLM-class models; ViT-Base is fine on Pi 5 (~340MB), Whisper-tiny is fine on Pi Zero 2 W (~75MB), Whisper-large is NOT (~3GB on a 512MB Zero). Recommend pinning the Whisper variant explicitly in the smoke-test acceptance ("Whisper-small or smaller for Pi 5; Whisper-tiny for Zero") — current text just says "Whisper" which leaves a silent-large-v3 footgun.
License audit gate must be acceptance criterion, not narrative — clinical deployment requires per-model auditing. Confirmed: present as a hardened checkbox in acceptance criteria.
Cross-surface invariants needing tightening:
- Whisper smoke spec drifts: "10-second German clinical clip" is not a pinned artefact; needs corpus identifier, SHA, reference transcript, numerical CER/BLEU threshold. As written, this is paper-passable.
- Pinned model versions: ONNX clinical BERT model + onnxruntime reference commit + whisper.cpp reference commit + ViT-Base weights need SHA-pinning analogous to E1's GGUF/llama.cpp pin policy. Currently named but not pinned.
- AMX-not-polyfill: E1 has this as an explicit acceptance criterion; E2 mentions "AMX flags for SPR" but does not gate on AMX-was-actually-exercised. Recommend lifting E1's wording into E2 acceptance for symmetry.
Tolerance propagation from D1: explicit and correct.
Phase ordering (E2 after E1): explicit and correct.
Worker: E2 (ensemble: model-frameworks track)
Surface: model frameworks
Filed in:
AdaWorldAPI/ndarraybecause ndarray is the obligatory spine;AdaWorldAPI/candledoes not exist yet — repo creation is part of execution under this plan, not a precondition.Why
After E1 confirms the spine pattern works in
burn-fork(GGUF end-to-end parity passes against the upstream reference), the same playbook applies to HuggingFacecandle. Today HuggingFacecandlelinksaccelerate-srcon macOS andcblason Linux. Acandle-forkthat swaps both forAdaWorldAPI/ndarrayunlocks four production surfaces we currently cannot serve coherently from one BLAS substrate:The strategic point: one ndarray, two deployment modes, hardware-accelerated on both. The same model binary runs on Railway SPR-AMX (cohort/batch path) and on a Pi 5 NEON box at the practice (edge inference). For a German Hausarzt that cannot ship patient images to cloud (GDPR + Krankenhaus-IT-Sicherheitsgesetz), local ViT screening on a Pi 5 is the difference between deployable and non-deployable. That asymmetry — same binary, two accelerator backends, both first-class — is the moat.
What
Create
AdaWorldAPI/candleas a fork ofhuggingface/candle, swap the BLAS dependency toAdaWorldAPI/ndarray, and validate four smoke surfaces plus an edge deployment matrix.Concrete items
AdaWorldAPI/candlefromhuggingface/candle. Tag the upstream commit at fork time; default branch tracks our integration branch, not upstreammain.accelerate-src(macOS) andcblas(Linux) withAdaWorldAPI/ndarrayas the BLAS substrate across the workspaceCargo.tomland per-crate manifests. Remove the platformcfggates that pick betweenaccelerate-srcandcblas; the ndarray dependency is the single source.git fetch upstream+ integration-branch rebase, conflict resolution log, regression-gate (smoke tests must pass before merging upstream into our default branch). This is tracked here but the recurring work happens in the candle-fork repo, not on this issue.candle-onnxcrate — wire ONNX gemm through ndarray (same gemm path the rest of candle now uses). No bypass back to cblas.medbert-deor equivalent), parity vsonnxruntimeon 10 sample sentences. Top-K embeddings agree within the tolerance defined by D1's parity harness.ViT-Base, parity vs reference on 10 images.whisper.cppreference.--no-default-features --features simd-neonfor Pi targets; AMX flags for SPR.Architecture
Spine pattern (repeated from E1). The fork's job is to retarget BLAS at the spine; everything above the BLAS line — model loading, kernels, control flow — is upstream code we keep merging in. The fork is small, structural, and audit-friendly; it does not own the model logic.
Edge moat. The same compiled artefact for ViT (or Whisper, or ONNX) targets
--features amxon SPR and--features simd-neonon Pi. ndarray decides which kernel to dispatch at runtime via the same gemm entry point. No model-side#[cfg], no separate Pi build of model code. That property is what makes "ship same binary to practice + cohort engine" feasible.Acceptance criteria
AdaWorldAPI/candlerepo exists, forked fromhuggingface/candle, fork commit tagged.accelerate-srcandcblasremoved across the workspace,AdaWorldAPI/ndarrayis the only BLAS dependency.x86_64-linuxandaarch64-linux.--no-default-features --features simd-neonfor Pi).MERGE_CADENCE.mdor equivalent).Out of scope
MedCare-rshandler wiring (ONNX/ViT/Whisper plumbed into clinical request paths) — a separate item, downstream of this one.onnxruntimefor training; this fork is inference-only.Dependencies