1.12.0 — Claude Code on a llama.cpp or vLLM server - #38
Merged
Merged
Conversation
alc could launch every agent but Claude Code on a llama-server. There was no llama.cpp kind, so a llama-server was set up as `vllm`, and vllm was a Responses-and-chat kind that Claude Code refused outright: "not compatible with claude". Both servers have answered Anthropic Messages at their root for some time; nothing in alc knew. `llamacpp` is now a provider kind of its own (`--kind llamacpp`, or llama.cpp, llama-cpp, llama-server; `--llamacpp` / `--llama-cpp` to select one): llama-server's default http://localhost:8080/v1, no auth unless a key is saved or LLAMA_API_KEY - the variable llama-server itself reads - is set, and every one of the eight agents. OpenCode takes its Chat Completions route, because the OpenAI SDK fails on llama-server's streamed Responses events; Codex takes Responses; a profile written with the bare root still reaches the OpenAI routes. A llama.cpp, vLLM or Ollama profile now runs Claude Code whatever OpenAI flavour it names for the other agents, so the chat-only `vllm` profiles people wrote for a llama-server work too. Claude Code gets the server's root rather than /v1, every model alias pinned to the one model it serves, and the context one request really gets: llama.cpp's per-slot n_ctx from /props, or vLLM's max_model_len, asked with the profile's key. Two things a real llama-server needed beyond the pins, both found by driving Claude Code 2.1.280 against Qwen3.8-27B on a shared llama.cpp b11100: - For a model it does not recognise, Claude Code sends its environment block as a `role: "system"` message after the first user turn. Qwen's chat template raises on any system message but the first, llama.cpp answers 500 "System message must be at the beginning", and Claude Code's own recovery only acts on a 400 in wording it knows, so it resent the request eleven times and gave up. For llama.cpp and vLLM, which render the model's own Jinja template, alc sets CLAUDE_CODE_MODEL_CAPABILITIES to turn mid-conversation system messages off; the block then rides in the first user turn as a <system-reminder>. Ollama renders its own templates and keeps its request shape. - llama-server sends its response headers at once and then nothing until the prompt is read - half a minute for 11k tokens there, a quarter of an hour for 200k. Claude Code's stream watchdogs end that silence after five minutes on any ANTHROPIC_BASE_URL, whatever API_FORCE_IDLE_TIMEOUT says, so every local server now also gets CLAUDE_STREAM_IDLE_TIMEOUT_MS at its thirty-minute ceiling. A value the user set is left alone, as before. `alc doctor` gains a llama.cpp and vLLM section: the server and its build, whether it takes the key, whether it lists the model, the context one request gets, and whether /v1/messages exists - asked with an empty body, which the server refuses before any model reads a token. A turned-off profile now says it is turned off instead of blaming its protocol. Verified live against that server, through this build: a Read tool call, and a Write plus a Bash run, each answered correctly in under a minute with no API error. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01Ld7aSDcdbdqqdefotHpmKq
The llamacpp kind joins the provider table, the shortcut flags and the examples, and vllm becomes Claude-ready. Local models now covers all three local servers: what alc pins, where the context comes from, and a section for llama.cpp and vLLM - the profile, the key, why mid-conversation system messages are turned off, why the stream watchdog goes to thirty minutes, and what alc doctor's new section checks. Troubleshooting gains the Jinja template's 500 and the turned-off profile, and the two print-mode notices now say they come with local models too. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01Ld7aSDcdbdqqdefotHpmKq
The zh-TW README and site pages follow the English: the llamacpp kind, vllm as Claude-ready, the llama.cpp and vLLM section of Local models, and the new troubleshooting entries. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01Ld7aSDcdbdqqdefotHpmKq
Claude Code on a llama.cpp or vLLM server. `llamacpp` is a provider kind of its own, reachable from all eight agents, and a llama.cpp, vLLM or Ollama profile now runs Claude Code whatever OpenAI protocol it names for the others - including the `vllm` profiles people had pointed at a llama-server. Claude Code gets the server's root, every model alias on the one model it serves, and the context one request really gets, read with the profile's key. Two fixes a real llama-server needed: Claude Code no longer sends a mid-conversation system message that the model's Jinja chat template refuses with a 500, and its stream watchdog waits out a long silent prompt read instead of cutting it off at five minutes. `alc doctor` checks llama.cpp and vLLM servers, and a turned-off profile says it is turned off. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01Ld7aSDcdbdqqdefotHpmKq
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
alc launched every agent but Claude Code on a llama-server. There was no llama.cpp kind, so people set one up as
vllm— andvllmwas a Responses-and-chat kind that Claude Code refused ("not compatible with claude"), even though llama.cpp and vLLM both answer Anthropic Messages at their root.llamacpp(--kind llamacpp/llama.cpp/llama-cpp/llama-server; shortcut--llamacpp, alias--llama-cpp): llama-server's defaulthttp://localhost:8080/v1, no auth unless a key is saved orLLAMA_API_KEY(the variable llama-server reads for--api-key) is set, all eight agents. OpenCode takes Chat Completions (the OpenAI SDK fails on llama-server's streamed Responses events), Codex takes Responses, and a root URL without/v1still reaches the OpenAI routes.vllmprofiles pointed at a llama-server work as they are. Claude Code gets the server's root (not/v1), every model alias pinned to the one model it serves, and the context one request gets: llama.cpp's per-slotn_ctxfrom/props, or vLLM'smax_model_len, read with the profile's key.role: "system"message after the first user turn. Qwen's Jinja template raises on it, llama.cpp answers500 System message must be at the beginning, and Claude Code's capability recovery only acts on a 400, so it retried 11 times and gave up. alc setsCLAUDE_CODE_MODEL_CAPABILITIES=-mid_conv_system,-mid_conv_tool_change(the documented*_SUPPORTED_CAPABILITIESvariables are ignored behindANTHROPIC_BASE_URL); the block then rides in the first user turn as a<system-reminder>. Ollama renders its own templates and keeps its request shape.ANTHROPIC_BASE_URLregardless ofAPI_FORCE_IDLE_TIMEOUT, so every local server also getsCLAUDE_STREAM_IDLE_TIMEOUT_MS=1800000(the ceiling). User-set values win, as before.alc doctorgains a llama.cpp and vLLM section: server and build, key accepted or not, model listed or not, per-request context, and whether/v1/messagesexists (asked with an empty body, refused before any model reads a token).provider 'vllm' is turned off; turn it on with alc config upsert vllm --enable --model <id>instead of blaming its protocol.chore: release 1.12.0.Test plan
/v1), llamacpp × all 8 agents, turned-off profile message, the full Claude document for a keyed and a keyless llama.cpp profile, capability switch only for llama.cpp/vLLM and never over a user-set value, OpenCode on Chat Completions, root URL normalised for Codex and Claude,/v1/modelsand/propsparsing (llama.cpp and vLLM shapes)--llamacpp --dry-run claudereads the per-slot context with the key and never prints it; dry run with no server;alc doctorsection; doctor names the missing key;--vllmon the starter template says turned offcargo test --all-targets(673 unit + 90 integration),cargo clippy --all-targets --all-features -- -D warningson stable 1.98.1,cargo fmt -- --check,npm run buildof the site (en + zh-TW, broken links/anchors throw)--api-key): aReadtool call and aWrite+Bashrun, each correct in under a minute with no API error;alc doctorreports the build, 262144 tokens and Anthropic Messages. Before the capability switch the same request failed 11/11 with the Jinja 500./v1/messagesand these same env vars for Claude Code)🤖 Generated with Claude Code
https://claude.ai/code/session_01Ld7aSDcdbdqqdefotHpmKq