Repository navigation
Tags: feiyunwill/llama.cpp
Tags
ggml: refactor selective expert copying to user code (ggml-org#29943) * ggml: refactor selective expert copying to user code * tests: enroll two models into selective expert copy test * tests: use deepseek2 as test model * improve comment in ggml-backend.h Co-authored-by: Georgi Gerganov <[email protected]> * cont: fix whitespace * cont : better comments Co-authored-by: Georgi Gerganov <[email protected]> --------- Co-authored-by: Georgi Gerganov <[email protected]>
hexagon: ssm-conv updates (ggml-org#29971) * hexagon: ssm-conv double-buffered DMA for prefill and decode restructuring * hex-ssm-conv: remove divs from loops and fix trace events * hex-dma: improved SSM_CONV dma pipeline and streamlined dma_queue --------- Co-authored-by: Max Krasnyansky <[email protected]>
metal : few-row MMA mat-mul (ggml-org#29869) * metal : few-row MMA mat-mul and batched copies for speculative decoding Speculative decoding verifies a few draft tokens per step. Without the tensor API, Metal ran these mat-muls with the mat-vec kernels, whose time grows with every src1 row, so DFlash2 decoding on an M3 Ultra was slower than serial decoding. - add mat-mul kernels for 2..16 src1 rows on 8x8 simdgroup matrices: each weight is dequantized once for all rows, and the simdgroups of a threadgroup split K. Q4_0, Q8_0 and Q5_K have their own kernels, F32, F16, Q4_1, Q5_0, Q5_1, Q4_K and Q6_K use a generic path over the 16-weight dequantizers, and Q4_0 at 2 rows uses a 2-row variant of the mat-vec kernel - use them only on MTLGPUFamilyApple7+ without the tensor API, from the row count at which they beat the mat-vec kernels on an M3 Ultra (F32: 6, F16, Q4_K, Q5_0, Q5_1: 3, other types: 2) - fusion table: MUL_MAT + ADD adds a same-shape residual in the MMA store, and up to 16 adjacent same-layout f32 copies between the same two tensors run as one dispatch - the fusion checks and ggml_graph_optimize take the device props, so the reorder packs MUL_MAT + ADD only on devices that can fuse it, at every src1 row count - views do not count toward GGML_METAL_FUSION_MAX when the reorder packs a group, so 16 recurrent state snapshot copies with views between them stay one group - the encoder checks the inner nodes of a fused group for concurrency, tracks written views by their extent, and does not count the destination of a CPY as a read - CONCAT splits long rows across threadgroups when there are few rows - tests: few-row MUL_MAT, MUL_MAT_ADD, CPY_BATCH and CONCAT cases in test-backend-ops (with a prepare_graph hook for the copy order), test-metal-graph-optimize, test-metal-cpy-batch-alias * metal : remove the CPY_BATCH fusion and the memory range changes Remove the batched copy fusion with its kernel and tests, and revert the memory range changes, as suggested in review. The memory ranges, the graph reorder and the CPY encoder are again the same as on master. * cont : clean-up * cont : drop has_tensor gate * cont : clean-up operand/residual logic * cont : drop Q4_0 ne11=2 special-case * cont : add kernels/mul_mv_mma.metal * cont : consolidate mma pipeline selection logic * cont : decouple fusion logic from device props --------- Co-authored-by: Georgi Gerganov <[email protected]>
model : Add LFM2.5-Encoder-350M and LFM2.5-Encoder-230M (ggml-org#29862) Register `Lfm2BidirectionalForMaskedLM` architecture for LFM2.5-Encoder models.
jinja : support coerced array attributes (ggml-org#29574) * support coerced array attributes * add tests
vulkan : fuse UNARY(GELU|SIGMOID|SILU|SOFTPLUS) + MUL (ggml-org#27220) * vulkan : fuse UNARY(SIGMOID|SILU|SOFTPLUS) + MUL * vulkan : fuse UNARY(SIGMOID|SILU|SOFTPLUS) + MUL - implement fusion in unary.comp behind UNARY_MUL_FUSION ifdef, specialized pipelines per op instead of runtime branching - fuse adjacent nodes only, ordering handled by graph_optimize - drop runtime consumer scan and pending_unary_mul deferral * vulkan : fuse UNARY(GELU|SIGMOID|SILU|SOFTPLUS) + MUL 1. GELU: gelu_mul_f32/f16 pipelines registered, CREATE_UNARY_MUL(gelu), GELU in dispatch + fuse gate + perf fusion name 2. Renamed/moved: gate is now ggml_vk_can_fuse_unary_mul(cgraph, unary_idx, mul_idx), placed with the other can-fuse helpers 3. norepeat both variants: each op gets plain (spec {0}) + _norepeat (spec {1}) pipelines from the same SPIR-V, selected via ggml_are_same_shape(src0, src1); the shape gate now allows broadcast (other dims equal-or-1) 4. graph_optimize: lambda deleted; standard "// UNARY + MUL: pull the consuming MUL forward" block added alongside the SSM_CONV/ROPE/MUL_MAT reorderings, with the same "other src must be weights or already processed" readiness check * vulkan : align unary_mul fusion with binary kernel layout, relax gelu test tolerance - schedule the fused kernel like mul.comp (256 threads x 2 unrolled iterations), recovering a 10-18% prompt-processing regression - allow 5e-7 f32 error for gelu_mul: the shader evaluates gelu with an exp-based tanh identity while the CPU reference uses tanhf (~1 ulp) * vulkan : use ggml_can_repeat in UNARY+MUL fusion shape check The fused kernel indexes src1 via per-dim fastmod (generic_binary_head.glsl), which is exact whenever the other operand tiles into the unary result -- not just when its dims are equal or 1. Replace the hand-rolled loop with ggml_can_repeat(other, unary) so the check matches the kernel's actual capability and reuses the standard helper. Argument order matters: reversed, it would wrongly admit graphs where the unary result is mul->src[1] and the other operand is larger, producing truncated output. Also add a rep_ne0 layout to the fused unary+mul backend tests covering a non-1 repeat factor along dim 0. * vulkan : fuse UNARY+MUL pairs separated by zero-compute nodes gemma4's per-layer embedding gating builds gelu -> view_2d_slice -> mul, where the intervening view is a zero-compute node aliasing an input that was computed much earlier. Strict adjacency requirements meant neither CUDA nor the vulkan unary+mul fusion handled this pattern. Extend ggml_vk_graph_optimize to detect a UNARY whose consuming MUL is separated only by unscheduled zero-compute nodes (GGML_OP_NONE, VIEW, RESHAPE, TRANSPOSE, PERMUTE) and schedule those nodes ahead of the pair, making it adjacent so the existing fusion applies. The reorder is guarded by ggml_vk_can_fuse_unary_mul, a source-availability check for every interleaved node, and the protected fusion patterns (topk_moe*, snake); if fusion is later rejected the reordered graph still executes correctly, just unfused. Add a view_mid layout to the fused unary+mul backend tests replicating the gemma4 pattern. * vulkan : support OP-on-B in UNARY+MUL fusion Some models apply the unary activation to the smaller MUL operand, e.g. qwen3next/qwen35moe shared-expert gating builds ffn_shexp * sigmoid(gate) with a [1,n_tokens] gate tensor. This shape was correctly rejected before: the fused kernel derives its iteration extent from the unary tensor and would leave most of the destination unwritten, and the generic same-shape requirement in ggml_can_fuse blocked the pair outright. Add UNARY_MUL_B_FUSION shader variants computing dst = src0 * OP(src1): the OP operand rides the existing per-dim fastmod indexing, while the iteration extent now comes from mul. Route {UNARY, MUL} pairs through a local can-fuse variant that drops the generic same-shape rule and instead requires the unary result to tile into mul->src[0] (ggml_can_repeat); pairs with the unary as src0 keep the previous direction check, and equal-shape pairs keep using the original pipelines. Add a "gate" layout to the fused unary+mul backend tests covering the shared-expert gate shape for gelu/sigmoid/silu/softplus in f32 and f16. * vulkan : fold unary+mul view-hoisting into graph_optimize dep checks Replace the dedicated UNARY + EMPTY* + MUL scanning block with two small extensions to the existing scheduling logic: - a consuming MUL may now join its in-set UNARY across a gap of unused zero-compute nodes (NONE/VIEW/RESHAPE/TRANSPOSE/PERMUTE), instead of requiring strict adjacency - while doing so, such zero-compute blockers are ignored for this pair Fusion validity is still decided later by ggml_vk_can_fuse at dispatch time, so a rejected pair simply executes adjacent-but-unfused. Note the relaxation must stay scoped to this pattern: exempting zero-compute blockers globally reproduces silent output corruption on gemma3n. * vulkan : select unary_mul OP-on-B via specialization constant Replace the UNARY_MUL_B_FUSION compile-time shader variants with an op_on_b specialization constant on the existing unary_mul SPIR-V, mirroring how the norepeat flag is handled. The four {op}_mul_b_{f32,f16} shader artifacts are gone - the OP-on-B pipelines reuse the base SPIR-V with two-entry {norepeat, op_on_b} spec lists - and the duplicated store expression is collapsed into a single runtime branch that the driver prunes per specialization. The constant is declared only under UNARY_MUL_FUSION so every other binary pipeline keeps its single-entry specialization list. * vulkan : replace unary_mul pipeline switches with a lookup table Collapse the four nested selection switches in ggml_vk_unary_mul into a single indexed lookup against a pipeline_unary_mul[4][2][2][2] table ([unary op][f16][norepeat][op_on_b]), whose trailing dims mirror the {norepeat, op_on_b} spec constant list. The op axis uses a small shared index helper that also replaces the switch in ggml_vk_can_fuse_unary_mul, making it the only place that maps ops to the table. Pipeline names are unchanged. Adding another supported op now requires one macro invocation line and one helper case instead of edits in four separate switches. * vulkan : use ggml_can_fuse_subgraph for unary_mul pairs Replace the hand-rolled pair validation in ggml_vk_can_fuse_unary_mul_pair (bounds, op match, compute flags, single-use elision) with the shared ggml_can_fuse_subgraph helper; backend-specific shape/type rules remain in ggml_vk_can_fuse_unary_mul. Unlike ggml_can_fuse, the subgraph helper has no same-shape requirement, so it covers both operand slots including OP-on-B gates, and additionally rejects intermediates flagged as graph outputs and validates view-source confinement. The outputs parameter takes absolute node indices into the cgraph. * Fix Whitespace * vulkan : drop redundant unary_mul gap check in graph_optimize The zero-compute nodes separating a UNARY from its consuming MUL are already scheduled ahead of the pair by pass 2 of an earlier optimization window, so the scoped gap tolerance added for this pattern is unreachable in practice - disabling it leaves gemma-3n dispatch counts unchanged (841 GELU_MUL per pass). Remove the flag, the empty blocker exemption, and the now-unused gap helper, restoring the strict adjacency requirement of the UNARY -> MUL pull-forward. Keep the relaxation scoped out entirely: generalizing "zero-compute nodes never block" beyond this pattern previously reproduced silent output corruption on gemma3n. * vulkan: fix whitespace (tab in indent) * vulkan: fix whitespace (extra blank line) * vulkan : move op_on_b spec constant to unary.comp op_on_b is only used by the fused unary*mul path. Keep generic_binary_head.glsl generic by defining it in unary.comp instead. Same constant_id=1 and guard, no functional change. * vulkan : make RMS_NORM/UNARY fusion gap-tolerant for views Strict j==c+1 blocked RMS_NORM->MUL and UNARY->MUL when a VIEW sits between (e.g. rms_norm -> view -> mul). Allow c==back() with an empty-or-scheduled gap, matching the review suggestion to check src linkage instead of adjacency. Scoped to the two blessed pairs; safe because gaps can only contain zero-compute nodes. * vulkan : trim comments in UNARY+MUL fusion Assisted-by: Muse Spark
model: add DSpark support for Nemotron3.5 (ggml-org#27804) * model: add DSpark support for Nemotron3.5 * Update src/models/dflash.cpp Co-authored-by: Sigbjørn Skjæret <[email protected]> --------- Co-authored-by: Sigbjørn Skjæret <[email protected]> Co-authored-by: Xuan Son Nguyen <[email protected]>
common: add system-level config file (ggml-org#26118) * common: Add CLI > ENV > models-presets > INI precedence 1. CLI flags have the highest precedence 2. ENV vars have the second-highest precedence 3. System and User configs have the lowest precedence - Linux/BSD/Mac - /etc/llama.cpp/config.ini < ${XDG_CONFIG_HOME:-~/.config}/llama.cpp/config.ini - Windows - %PROGRAMDATA%\llama.cpp\config.ini < %APPDATA%\llama.cpp\config.ini * fix UB * use common_get_env * ignore_unknown_keys * nits * add docs --------- Co-authored-by: Xuan Son Nguyen <[email protected]>
ggml-webgpu: improve i-quants mul_mat performance and speed up prefill ( ggml-org#24530) * Improve prefill speeds for i-quants * Fix #if defined() usage in preprocessor guards.
ggml-webgpu: Add clang-format job (ggml-org#24308) * Add clang-format job * try local formatting
PreviousNext