feat(gateway): accept Gemini clients, and the Responses API over a WebSocket - #43
Merged
Merged
Conversation
…bSocket
Gemini: POST /v1beta/models/{model}:generateContent and
:streamGenerateContent (also under /v1), through the same pipeline as
the other surfaces. The model and the stream flag come from the path.
Inside, a Gemini stream is always SSE; a caller that did not ask for
alt=sse gets it reframed as one JSON array at the end. A Gemini upstream
gets the request as sent, at the routed model's path. Errors are in
Gemini's shape. GET /v1beta/models lists models in Gemini's shape, and
the model name in modelVersion is the caller's, like elsewhere.
The API key is read from x-api-key, x-goog-api-key or ?key= as well as
Authorization: Bearer, so Anthropic and Gemini SDKs work with their own
settings. The query never reaches an upstream.
Responses over WebSocket: GET /v1/responses upgraded. Each
response.create frame is one request through generate(), so it is
limited, routed and converted, inspected, billed and logged exactly like
a streamed POST /v1/responses; its SSE events go back as text frames.
A refusal is a response.failed frame and the connection stays open.
Co-Authored-By: Claude Opus 5.5 <[email protected]>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Brings two client surfaces the desktop gateway already has to this side, both through the existing pipeline rather than beside it.
Gemini-format clients
POST /v1beta/models/{model}:generateContentand:streamGenerateContent(also under/v1/models/…). The model and the stream flag come from the path; everything else — limits, budget, model access, content filter, hidden text, PII, cache (off, as for the other non-Chat surfaces), routing with failover, conversion to the route's format, tool-call inspection, output guardrails, billing, the audit row — is the samegeneratethe other three surfaces run.alt=sse, and the shaper, the inspection, the sniffer and the error frames all read SSE. A caller that did not passalt=ssegets the stream reframed as Gemini's JSON array as the last step (shaper::JsonArrayFramer), withContent-Type: application/json./v1beta/models/{routed model}:{action}(nomodelfield inserted, fields the IR cannot carry survive), like same-format forwarding for the other formats.modelVersioncarries the caller's model name, likemodeldoes elsewhere. Errors usetw_dialect::convert::error_body/error_framein Gemini's shape.GET /v1beta/modelslists the routed models in Gemini's shape.:countTokens,:embedContent) are refused with 400: they have no counterpart in another format.API key positions
The gateway middleware now reads the key from
x-api-key(Anthropic SDKs),x-goog-api-keyor?key=(Gemini SDKs) as well asAuthorization: Bearer, headers first, same order as the desktop. The query never reaches an upstream: outbound requests use the path and query the gateway builds. Request logging records paths only.Responses API over WebSocket
GET /v1/responseswithUpgrade: websocket, authenticated by the same middleware on the upgrade. Each{"type":"response.create", …}text frame (fields at the top level, or nested underresponse) is one request throughgenerateas a streamedPOST /v1/responses; its SSE events go back one JSON object per text frame. So every turn is rate-limited, budget-checked, routed and converted to any upstream format, inspected, billed and gets its own audit row — a socket-to-socket pipe would skip all of that, which is what the desktop'sws.rswarns about. Turns on a connection run in order; frames arriving mid-turn are queued; a client that leaves mid-turn cancels it (recorded like a dropped HTTP stream). A refusal or an unrecognised frame is aresponse.failedframe and the connection stays open.Not carried over from OpenAI's socket mode: a connection-local store of the previous response. Each turn goes upstream as its own request, so
previous_response_idworks only against an upstream that stored that response.Desktop-only concerns not copied: the upstream-socket dialing, proxy refusal, and per-connection secret ledger.
Docs
README endpoint lists and the console guide's endpoint card (en/zh).
Tests
modelVersionrewrite,response.createparsing, key positions.gateway_gemini.rs: Gemini client → OpenAI upstream answered in Gemini format and billed (ClickHouse row), SSE stream with?key=, JSON-array stream, Gemini → Gemini forwarded as sent (path,alt=sse, no key in the query, unknown fields kept), Gemini error shapes, missing key, Gemini model listing.gateway_responses_ws.rs: two turns on one connection answered, converted to Chat upstream, and each billed/logged; upgrade without a key refused with 401; a per-user request limit applying per turn withresponse.failed/rate_limit_exceededand the connection surviving it.--all-targets,--lib), 687 unit tests, i18n check, full integration suite on own containers (all new tests pass; the one failure,streaming_cache_hit_replays_assembled_sse, is a pre-existing race fixed separately).🤖 Generated with Claude Code