fix: bill a Chat stream whose caller did not ask for usage - #34
Merged
Merged
Conversation
A Chat stream forwarded as sent reached the upstream without `stream_options.include_usage` unless the caller had set it. The upstream then reports no usage, and the request was recorded as zero tokens: no quota, no budget debit, no cost. 1.0.2 estimated the count in that case; the estimate went with the old pipeline in #26, and nothing took its place. A Chat stream now always asks the upstream for its usage. When the caller did not, the shaper takes it back out of what the caller receives: the trailing usage-only chunk, and the `"usage": null` the upstream adds to every other chunk once asked. The sniffer reads the upstream's own bytes before the shaper, so billing sees the real count. Converted streams already asked for usage and write the caller's chunk only when the caller wanted it. Co-Authored-By: Claude Opus 5.5 <[email protected]>
4 tasks
fylorn
added a commit
that referenced
this pull request
Sep 24, 2026
… carried signatures Cached input was billed at the full input price. The prompt count put cache reads and writes together with plain input, so an Anthropic turn that read 10k tokens from its cache cost as much as sending them fresh, and weighed as much against rate limits and budgets. Models get three optional weights against the input baseline: cache read, cache write, and 1-hour cache write. Unset, they follow the input weight at Anthropic's ratios (0.1x, 1.25x, 2x). Cost and weighted tokens now price each bucket on its own weight. The audit row keeps the whole input in `input_tokens` and gains the cache split in its detail. The admin API and the model editor carry the new weights. A request the upstream reported no usage for was billed as zero tokens. #34 made Chat streams ask for usage, but usage can still be missing: the upstream ignores `stream_options`, or the caller leaves mid-stream and takes the final usage chunk with it. The count is now estimated at about four bytes a token from the request and from the answer that arrived, and the row says `usage_estimated`. A stream cut short keeps the upstream's input count and bills the larger of its running output count and the estimate. No answer at all still bills nothing. A request forwarded as sent now drops the `tw1.` reasoning signatures an earlier conversion wrote, via `tw_dialect::convert::strip_carried`, as the desktop gateway does. Anthropic refuses a whole request over them. Co-Authored-By: Claude Opus 5.5 <[email protected]>
fylorn
added a commit
that referenced
this pull request
Sep 24, 2026
… carried signatures (#40) Cached input was billed at the full input price. The prompt count put cache reads and writes together with plain input, so an Anthropic turn that read 10k tokens from its cache cost as much as sending them fresh, and weighed as much against rate limits and budgets. Models get three optional weights against the input baseline: cache read, cache write, and 1-hour cache write. Unset, they follow the input weight at Anthropic's ratios (0.1x, 1.25x, 2x). Cost and weighted tokens now price each bucket on its own weight. The audit row keeps the whole input in `input_tokens` and gains the cache split in its detail. The admin API and the model editor carry the new weights. A request the upstream reported no usage for was billed as zero tokens. #34 made Chat streams ask for usage, but usage can still be missing: the upstream ignores `stream_options`, or the caller leaves mid-stream and takes the final usage chunk with it. The count is now estimated at about four bytes a token from the request and from the answer that arrived, and the row says `usage_estimated`. A stream cut short keeps the upstream's input count and bills the larger of its running output count and the estimate. No answer at all still bills nothing. A request forwarded as sent now drops the `tw1.` reasoning signatures an earlier conversion wrote, via `tw_dialect::convert::strip_carried`, as the desktop gateway does. Anthropic refuses a whole request over them. Co-authored-by: Claude Opus 5.5 <[email protected]>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Found while preparing 1.1.0. This is a regression from #26.
A Chat stream forwarded as sent reached the upstream without
stream_options.include_usageunless the caller had set it. Many clients don't set it. The upstream then reports no usage, and the request was recorded as zero tokens: no quota, no budget debit, no cost. 1.0.2 estimated the count in that case. The estimate went away with the old pipeline in #26, and nothing replaced it."usage": nullthe upstream adds to every other chunk once asked.Test plan
Unit: the shaper strips the usage chunk and
usage: nullwhen hiding, and leaves both alone when the caller asked.Integration
a_chat_stream_is_billed_when_the_caller_did_not_ask_for_usage:usage;gateway_logsrecords 5/2 tokens.It fails without the fix: the upstream gets a request without
include_usage.fmt, clippy
-D warnings, gateway unit tests (134)🤖 Generated with Claude Code