Skip to content

feat(gateway): accept Gemini clients, and the Responses API over a WebSocket - #43

Merged
fylorn merged 1 commit into
devfrom
feat/gemini-clients-and-responses-ws
Sep 24, 2026
Merged

fylorn merged 1 commit into
devfrom
feat/gemini-clients-and-responses-ws

Conversation

@fylorn

@fylorn fylorn commented Sep 24, 2026

Copy link
Copy Markdown
Contributor

Brings two client surfaces the desktop gateway already has to this side, both through the existing pipeline rather than beside it.

Gemini-format clients

  • POST /v1beta/models/{model}:generateContent and :streamGenerateContent (also under /v1/models/…). The model and the stream flag come from the path; everything else — limits, budget, model access, content filter, hidden text, PII, cache (off, as for the other non-Chat surfaces), routing with failover, conversion to the route's format, tool-call inspection, output guardrails, billing, the audit row — is the same generate the other three surfaces run.
  • Inside, a Gemini stream is always SSE. Upstreams are asked for alt=sse, and the shaper, the inspection, the sniffer and the error frames all read SSE. A caller that did not pass alt=sse gets the stream reframed as Gemini's JSON array as the last step (shaper::JsonArrayFramer), with Content-Type: application/json.
  • A Gemini upstream gets the request as sent at /v1beta/models/{routed model}:{action} (no model field inserted, fields the IR cannot carry survive), like same-format forwarding for the other formats.
  • modelVersion carries the caller's model name, like model does elsewhere. Errors use tw_dialect::convert::error_body/error_frame in Gemini's shape.
  • GET /v1beta/models lists the routed models in Gemini's shape.
  • Actions other than the two generations (:countTokens, :embedContent) are refused with 400: they have no counterpart in another format.

API key positions

The gateway middleware now reads the key from x-api-key (Anthropic SDKs), x-goog-api-key or ?key= (Gemini SDKs) as well as Authorization: Bearer, headers first, same order as the desktop. The query never reaches an upstream: outbound requests use the path and query the gateway builds. Request logging records paths only.

Responses API over WebSocket

GET /v1/responses with Upgrade: websocket, authenticated by the same middleware on the upgrade. Each {"type":"response.create", …} text frame (fields at the top level, or nested under response) is one request through generate as a streamed POST /v1/responses; its SSE events go back one JSON object per text frame. So every turn is rate-limited, budget-checked, routed and converted to any upstream format, inspected, billed and gets its own audit row — a socket-to-socket pipe would skip all of that, which is what the desktop's ws.rs warns about. Turns on a connection run in order; frames arriving mid-turn are queued; a client that leaves mid-turn cancels it (recorded like a dropped HTTP stream). A refusal or an unrecognised frame is a response.failed frame and the connection stays open.

Not carried over from OpenAI's socket mode: a connection-local store of the previous response. Each turn goes upstream as its own request, so previous_response_id works only against an upstream that stored that response.

Desktop-only concerns not copied: the upstream-socket dialing, proxy refusal, and per-connection secret ledger.

Docs

README endpoint lists and the console guide's endpoint card (en/zh).

Tests

  • Unit: Gemini path parsing, JSON-array reframing, modelVersion rewrite, response.create parsing, key positions.
  • Integration gateway_gemini.rs: Gemini client → OpenAI upstream answered in Gemini format and billed (ClickHouse row), SSE stream with ?key=, JSON-array stream, Gemini → Gemini forwarded as sent (path, alt=sse, no key in the query, unknown fields kept), Gemini error shapes, missing key, Gemini model listing.
  • Integration gateway_responses_ws.rs: two turns on one connection answered, converted to Chat upstream, and each billed/logged; upgrade without a key refused with 401; a per-user request limit applying per turn with response.failed/rate_limit_exceeded and the connection surviving it.
  • Local: fmt, clippy (--all-targets, --lib), 687 unit tests, i18n check, full integration suite on own containers (all new tests pass; the one failure, streaming_cache_hit_replays_assembled_sse, is a pre-existing race fixed separately).

🤖 Generated with Claude Code

…bSocket

Gemini: POST /v1beta/models/{model}:generateContent and
:streamGenerateContent (also under /v1), through the same pipeline as
the other surfaces. The model and the stream flag come from the path.
Inside, a Gemini stream is always SSE; a caller that did not ask for
alt=sse gets it reframed as one JSON array at the end. A Gemini upstream
gets the request as sent, at the routed model's path. Errors are in
Gemini's shape. GET /v1beta/models lists models in Gemini's shape, and
the model name in modelVersion is the caller's, like elsewhere.

The API key is read from x-api-key, x-goog-api-key or ?key= as well as
Authorization: Bearer, so Anthropic and Gemini SDKs work with their own
settings. The query never reaches an upstream.

Responses over WebSocket: GET /v1/responses upgraded. Each
response.create frame is one request through generate(), so it is
limited, routed and converted, inspected, billed and logged exactly like
a streamed POST /v1/responses; its SSE events go back as text frames.
A refusal is a response.failed frame and the connection stays open.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@fylorn
fylorn merged commit ed1c8b5 into dev Sep 24, 2026
6 checks passed
@fylorn
fylorn deleted the feat/gemini-clients-and-responses-ws branch September 24, 2026 08:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant