Skip to content

1.12.0 — Claude Code on a llama.cpp or vLLM server - #38

Merged
treeleaves30760 merged 4 commits into
mainfrom
feat/llamacpp-claude
Sep 23, 2026
Merged

treeleaves30760 merged 4 commits into
mainfrom
feat/llamacpp-claude

Conversation

@treeleaves30760

Copy link
Copy Markdown
Owner

Summary

alc launched every agent but Claude Code on a llama-server. There was no llama.cpp kind, so people set one up as vllm — and vllm was a Responses-and-chat kind that Claude Code refused ("not compatible with claude"), even though llama.cpp and vLLM both answer Anthropic Messages at their root.

  • New kind llamacpp (--kind llamacpp / llama.cpp / llama-cpp / llama-server; shortcut --llamacpp, alias --llama-cpp): llama-server's default http://localhost:8080/v1, no auth unless a key is saved or LLAMA_API_KEY (the variable llama-server reads for --api-key) is set, all eight agents. OpenCode takes Chat Completions (the OpenAI SDK fails on llama-server's streamed Responses events), Codex takes Responses, and a root URL without /v1 still reaches the OpenAI routes.
  • Claude Code on any local server. A llama.cpp, vLLM or Ollama profile now runs Claude Code whatever OpenAI protocol it names for the other agents, so existing chat-only vllm profiles pointed at a llama-server work as they are. Claude Code gets the server's root (not /v1), every model alias pinned to the one model it serves, and the context one request gets: llama.cpp's per-slot n_ctx from /props, or vLLM's max_model_len, read with the profile's key.
  • Mid-conversation system messages off for llama.cpp and vLLM. For a model it does not recognise, Claude Code 2.1.280 sends its environment block as a role: "system" message after the first user turn. Qwen's Jinja template raises on it, llama.cpp answers 500 System message must be at the beginning, and Claude Code's capability recovery only acts on a 400, so it retried 11 times and gave up. alc sets CLAUDE_CODE_MODEL_CAPABILITIES=-mid_conv_system,-mid_conv_tool_change (the documented *_SUPPORTED_CAPABILITIES variables are ignored behind ANTHROPIC_BASE_URL); the block then rides in the first user turn as a <system-reminder>. Ollama renders its own templates and keeps its request shape.
  • Stream watchdog. llama-server sends headers at once and then nothing until the prompt is read (≈30 s for 11k tokens on a shared 27B, minutes for long prompts). Claude Code's stream watchdogs end that after 5 min on any ANTHROPIC_BASE_URL regardless of API_FORCE_IDLE_TIMEOUT, so every local server also gets CLAUDE_STREAM_IDLE_TIMEOUT_MS=1800000 (the ceiling). User-set values win, as before.
  • alc doctor gains a llama.cpp and vLLM section: server and build, key accepted or not, model listed or not, per-request context, and whether /v1/messages exists (asked with an empty body, refused before any model reads a token).
  • A turned-off profile now says provider 'vllm' is turned off; turn it on with alc config upsert vllm --enable --model <id> instead of blaming its protocol.
  • Docs, EN and zh-TW: README, the zh-TW README, and the site pages (local models, providers, getting started, agents, usage, intro, troubleshooting).
  • chore: release 1.12.0.

Test plan

  • Unit: kind parsing/defaults, local servers speak Anthropic at their root whatever the protocol (llama.cpp, vLLM chat/responses, Ollama with and without /v1), llamacpp × all 8 agents, turned-off profile message, the full Claude document for a keyed and a keyless llama.cpp profile, capability switch only for llama.cpp/vLLM and never over a user-set value, OpenCode on Chat Completions, root URL normalised for Codex and Claude, /v1/models and /props parsing (llama.cpp and vLLM shapes)
  • Integration (fake llama-server that demands a key): --llamacpp --dry-run claude reads the per-slot context with the key and never prints it; dry run with no server; alc doctor section; doctor names the missing key; --vllm on the starter template says turned off
  • cargo test --all-targets (673 unit + 90 integration), cargo clippy --all-targets --all-features -- -D warnings on stable 1.98.1, cargo fmt -- --check, npm run build of the site (en + zh-TW, broken links/anchors throw)
  • Live, Claude Code 2.1.280 through this build against a shared lab llama-server (b11100, Qwen3.8-27B, --api-key): a Read tool call and a Write + Bash run, each correct in under a minute with no API error; alc doctor reports the build, 262144 tokens and Anthropic Messages. Before the capability switch the same request failed 11/11 with the Jinja 500.
  • Not tested live: vLLM (it documents /v1/messages and these same env vars for Claude Code)

🤖 Generated with Claude Code

https://claude.ai/code/session_01Ld7aSDcdbdqqdefotHpmKq

treeleaves30760 and others added 4 commits September 23, 2026 13:11
alc could launch every agent but Claude Code on a llama-server. There was
no llama.cpp kind, so a llama-server was set up as `vllm`, and vllm was a
Responses-and-chat kind that Claude Code refused outright: "not compatible
with claude". Both servers have answered Anthropic Messages at their root
for some time; nothing in alc knew.

`llamacpp` is now a provider kind of its own (`--kind llamacpp`, or
llama.cpp, llama-cpp, llama-server; `--llamacpp` / `--llama-cpp` to
select one): llama-server's default http://localhost:8080/v1, no auth
unless a key is saved or LLAMA_API_KEY - the variable llama-server itself
reads - is set, and every one of the eight agents. OpenCode takes its Chat
Completions route, because the OpenAI SDK fails on llama-server's streamed
Responses events; Codex takes Responses; a profile written with the bare
root still reaches the OpenAI routes.

A llama.cpp, vLLM or Ollama profile now runs Claude Code whatever OpenAI
flavour it names for the other agents, so the chat-only `vllm` profiles
people wrote for a llama-server work too. Claude Code gets the server's
root rather than /v1, every model alias pinned to the one model it serves,
and the context one request really gets: llama.cpp's per-slot n_ctx from
/props, or vLLM's max_model_len, asked with the profile's key.

Two things a real llama-server needed beyond the pins, both found by
driving Claude Code 2.1.280 against Qwen3.8-27B on a shared llama.cpp
b11100:

- For a model it does not recognise, Claude Code sends its environment
  block as a `role: "system"` message after the first user turn. Qwen's
  chat template raises on any system message but the first, llama.cpp
  answers 500 "System message must be at the beginning", and Claude Code's
  own recovery only acts on a 400 in wording it knows, so it resent the
  request eleven times and gave up. For llama.cpp and vLLM, which render
  the model's own Jinja template, alc sets CLAUDE_CODE_MODEL_CAPABILITIES
  to turn mid-conversation system messages off; the block then rides in
  the first user turn as a <system-reminder>. Ollama renders its own
  templates and keeps its request shape.
- llama-server sends its response headers at once and then nothing until
  the prompt is read - half a minute for 11k tokens there, a quarter of an
  hour for 200k. Claude Code's stream watchdogs end that silence after five
  minutes on any ANTHROPIC_BASE_URL, whatever API_FORCE_IDLE_TIMEOUT says,
  so every local server now also gets CLAUDE_STREAM_IDLE_TIMEOUT_MS at its
  thirty-minute ceiling. A value the user set is left alone, as before.

`alc doctor` gains a llama.cpp and vLLM section: the server and its build,
whether it takes the key, whether it lists the model, the context one
request gets, and whether /v1/messages exists - asked with an empty body,
which the server refuses before any model reads a token. A turned-off
profile now says it is turned off instead of blaming its protocol.

Verified live against that server, through this build: a Read tool call,
and a Write plus a Bash run, each answered correctly in under a minute
with no API error.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01Ld7aSDcdbdqqdefotHpmKq
The llamacpp kind joins the provider table, the shortcut flags and the
examples, and vllm becomes Claude-ready. Local models now covers all three
local servers: what alc pins, where the context comes from, and a section
for llama.cpp and vLLM - the profile, the key, why mid-conversation system
messages are turned off, why the stream watchdog goes to thirty minutes,
and what alc doctor's new section checks. Troubleshooting gains the Jinja
template's 500 and the turned-off profile, and the two print-mode notices
now say they come with local models too.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01Ld7aSDcdbdqqdefotHpmKq
The zh-TW README and site pages follow the English: the llamacpp kind,
vllm as Claude-ready, the llama.cpp and vLLM section of Local models, and
the new troubleshooting entries.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01Ld7aSDcdbdqqdefotHpmKq
Claude Code on a llama.cpp or vLLM server.

`llamacpp` is a provider kind of its own, reachable from all eight agents,
and a llama.cpp, vLLM or Ollama profile now runs Claude Code whatever
OpenAI protocol it names for the others - including the `vllm` profiles
people had pointed at a llama-server. Claude Code gets the server's root,
every model alias on the one model it serves, and the context one request
really gets, read with the profile's key.

Two fixes a real llama-server needed: Claude Code no longer sends a
mid-conversation system message that the model's Jinja chat template
refuses with a 500, and its stream watchdog waits out a long silent
prompt read instead of cutting it off at five minutes. `alc doctor` checks
llama.cpp and vLLM servers, and a turned-off profile says it is turned off.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01Ld7aSDcdbdqqdefotHpmKq
@treeleaves30760
treeleaves30760 merged commit 7955cbe into main Sep 23, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant