Skip to content

fix: bill a Chat stream whose caller did not ask for usage - #34

Merged
fylorn merged 1 commit into
devfrom
fix/chat-stream-usage
Sep 24, 2026
Merged

fylorn merged 1 commit into
devfrom
fix/chat-stream-usage

Conversation

@fylorn

@fylorn fylorn commented Sep 24, 2026

Copy link
Copy Markdown
Contributor

Summary

Found while preparing 1.1.0. This is a regression from #26.

A Chat stream forwarded as sent reached the upstream without stream_options.include_usage unless the caller had set it. Many clients don't set it. The upstream then reports no usage, and the request was recorded as zero tokens: no quota, no budget debit, no cost. 1.0.2 estimated the count in that case. The estimate went away with the old pipeline in #26, and nothing replaced it.

  • A Chat stream now always asks the upstream for its usage.
  • When the caller did not ask for it, the stream shaper takes it back out of what the caller receives:
    • the trailing usage-only chunk;
    • the "usage": null the upstream adds to every other chunk once asked.
  • The sniffer reads the upstream's own bytes before the shaper, so billing sees the real count.
  • Converted streams already asked for usage, and write the caller's usage chunk only when the caller asked for it.

Test plan

  • Unit: the shaper strips the usage chunk and usage: null when hiding, and leaves both alone when the caller asked.

  • Integration a_chat_stream_is_billed_when_the_caller_did_not_ask_for_usage:

    • the mock answers only a request that asks for usage;
    • the caller's stream carries no usage;
    • gateway_logs records 5/2 tokens.

    It fails without the fix: the upstream gets a request without include_usage.

  • fmt, clippy -D warnings, gateway unit tests (134)

🤖 Generated with Claude Code

A Chat stream forwarded as sent reached the upstream without
`stream_options.include_usage` unless the caller had set it. The
upstream then reports no usage, and the request was recorded as zero
tokens: no quota, no budget debit, no cost. 1.0.2 estimated the count
in that case; the estimate went with the old pipeline in #26, and
nothing took its place.

A Chat stream now always asks the upstream for its usage. When the
caller did not, the shaper takes it back out of what the caller
receives: the trailing usage-only chunk, and the `"usage": null` the
upstream adds to every other chunk once asked. The sniffer reads the
upstream's own bytes before the shaper, so billing sees the real
count. Converted streams already asked for usage and write the
caller's chunk only when the caller wanted it.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@fylorn
fylorn merged commit 9f46116 into dev Sep 24, 2026
@fylorn
fylorn deleted the fix/chat-stream-usage branch September 24, 2026 05:45
fylorn added a commit that referenced this pull request Sep 24, 2026
… carried signatures

Cached input was billed at the full input price. The prompt count put
cache reads and writes together with plain input, so an Anthropic turn
that read 10k tokens from its cache cost as much as sending them fresh,
and weighed as much against rate limits and budgets. Models get three
optional weights against the input baseline: cache read, cache write,
and 1-hour cache write. Unset, they follow the input weight at
Anthropic's ratios (0.1x, 1.25x, 2x). Cost and weighted tokens now price
each bucket on its own weight. The audit row keeps the whole input in
`input_tokens` and gains the cache split in its detail. The admin API
and the model editor carry the new weights.

A request the upstream reported no usage for was billed as zero tokens.
#34 made Chat streams ask for usage, but usage can still be missing: the
upstream ignores `stream_options`, or the caller leaves mid-stream and
takes the final usage chunk with it. The count is now estimated at about
four bytes a token from the request and from the answer that arrived,
and the row says `usage_estimated`. A stream cut short keeps the
upstream's input count and bills the larger of its running output count
and the estimate. No answer at all still bills nothing.

A request forwarded as sent now drops the `tw1.` reasoning signatures an
earlier conversion wrote, via `tw_dialect::convert::strip_carried`, as
the desktop gateway does. Anthropic refuses a whole request over them.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
fylorn added a commit that referenced this pull request Sep 24, 2026
… carried signatures (#40)

Cached input was billed at the full input price. The prompt count put
cache reads and writes together with plain input, so an Anthropic turn
that read 10k tokens from its cache cost as much as sending them fresh,
and weighed as much against rate limits and budgets. Models get three
optional weights against the input baseline: cache read, cache write,
and 1-hour cache write. Unset, they follow the input weight at
Anthropic's ratios (0.1x, 1.25x, 2x). Cost and weighted tokens now price
each bucket on its own weight. The audit row keeps the whole input in
`input_tokens` and gains the cache split in its detail. The admin API
and the model editor carry the new weights.

A request the upstream reported no usage for was billed as zero tokens.
#34 made Chat streams ask for usage, but usage can still be missing: the
upstream ignores `stream_options`, or the caller leaves mid-stream and
takes the final usage chunk with it. The count is now estimated at about
four bytes a token from the request and from the answer that arrived,
and the row says `usage_estimated`. A stream cut short keeps the
upstream's input count and bills the larger of its running output count
and the estimate. No answer at all still bills nothing.

A request forwarded as sent now drops the `tw1.` reasoning signatures an
earlier conversion wrote, via `tw_dialect::convert::strip_carried`, as
the desktop gateway does. Anthropic refuses a whole request over them.

Co-authored-by: Claude Opus 5.5 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant