Skip to content

fix(responses): recount tokens when a streamed response completes without usage - #41337

Merged
kerry-berri merged 10 commits into
mainfrom
litellm_fix_responses_stream_absent_usage_recount
Sep 16, 2026
Merged

kerry-berri merged 10 commits into
mainfrom
litellm_fix_responses_stream_absent_usage_recount

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • A streamed /v1/responses call whose completed event has no usage is billed $0
  • Every other streaming surface recounts tokens proxy-side; Responses did not

How it solves it:

  • On response.completed / response.incomplete with usage: null, estimate usage from the request input and the streamed output
  • Input is converted to chat messages first, so multimodal and structured input counts like the model sees it, not as JSON text
  • Output counts text deltas plus function-call, custom-tool and MCP argument deltas, so tool-only responses are not billed as empty
  • The estimate is stamped on the response before cost is computed, so the spend row is non-zero
  • If the estimate itself fails, the stream still completes and usage stays unset, as before
  • The terminal event's response may be a typed ResponsesAPIResponse or, when the OpenAI transformation falls back to model_construct, a plain dict; a dict is upgraded to a typed object and the terminal event is re-issued as a copy carrying it, so the estimate, cost stamping and logging all run on it instead of crashing or billing $0. Nothing is mutated in place, and the response cache write tolerates an upgraded object that cannot be serialized instead of failing the stream

User Flow

Before: a developer streaming from a Responses-compatible upstream that omits usage sees the call logged for free

  1. They send POST https://litellm-domain/v1/responses with "model": "gpt-5.6", "stream": true
  2. Output deltas arrive, then response.completed with "usage": null
  3. https://litellm-domain/ui/?page=logs shows the request with 0 tokens and $0 spend

After: the same stream is billed from a token recount

  1. They send the same POST https://litellm-domain/v1/responses with "stream": true
  2. Output deltas arrive, then response.completed now carries a usage object with estimated input and output tokens
  3. https://litellm-domain/ui/?page=logs shows the request with those token counts and non-zero spend

Relevant issues

Found by the scripted-provider cost suite in #41328

Affected release

Linear ticket

Resolves LIT-7840

Pre-Submission checklist

  • I have added meaningful tests
  • The handful of test files covering my change pass locally: uv run pytest tests/test_litellm/responses/test_streaming_iterator.py tests/test_litellm/responses/test_responses_utils.py plus the whole tests/test_litellm/responses directory (711 passed)
  • My PR passes all required CI/CD checks
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5

Screenshots / Proof of Fix

Regression tests stream output deltas followed by a response.completed event with usage: null and assert the completed response carries input and output token counts greater than zero with total = input + output; the same for a function-call-only stream; a multimodal input counts fewer tokens than its JSON serialization; custom-tool and MCP argument deltas are counted like function-call arguments; a completed event that does carry usage is left untouched; a model_constructed completed event whose response is a plain dict ends up typed with estimated usage and a stamped cost, with the cost calculator called on it and the yielded terminal event being the re-issued copy; a typed response that already has usage is passed through as the same object while one without usage gets a copy and the original is left untouched; an upgraded response whose JSON dump raises is skipped by the cache write without failing the stream; and a malformed input that makes the estimate raise leaves usage unset without failing the stream, with no monkeypatching involved. Live before/after against the scripted upstream will follow once #41328 lands, where the responses stream_no_usage case flips from expect_zero_bill to a non-zero bill

Type

🐛 Bug Fix

Caveats (if any)

Medium

  • The recount is an estimate from the tokenizer, like the chat completions fallback; provider usage is always preferred when present
  • response.failed is left alone and stays $0
  • An estimate that raises is logged at debug and the call stays at $0 rather than failing

Low


Note

Medium Risk
Changes proxy-side token estimation and spend calculation for Responses streaming when providers omit usage; incorrect estimates could misbill, though provider usage still wins and estimate failures fall back to $0 without failing the stream.

Overview
Fixes $0 billing when a streamed /v1/responses call finishes with usage: null on response.completed or response.incomplete.

The streaming iterator now accumulates output text plus function-call, custom-tool, and MCP argument deltas, then on terminal events estimates ResponseAPIUsage via token_counter (request input converted to chat messages first, so multimodal input is not counted as raw JSON). Provider-reported usage is unchanged; response.failed is not estimated. Estimates are attached through _billed_terminal_response, which copies the terminal event when needed so cost stamping and logging see typed usage without mutating the original object; plain-dict responses from model_construct are upgraded the same way.

Cache persistence uses _dump_json_safely so an unserializable completed response skips the cache write instead of breaking the stream. Regression tests cover text-only and tool-only streams, multimodal input, estimate failures, dict responses, and cache serialization edge cases.

Reviewed by Cursor Bugbot for commit ff878e7. Bugbot is set up for automated code reviews on this repo. Configure here.

Link to Devin session: https://app.devin.ai/sessions/25b6f01a2b594d44a5e2488dc54d010c
Open in Devin Desktop: https://app.devin.ai/desktop/session/25b6f01a2b594d44a5e2488dc54d010c?variant=devin

…t omits usage

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration
devin-ai-integration Bot requested a review from a team September 16, 2026 00:43
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@codspeed

codspeed Bot commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_fix_responses_stream_absent_usage_recount (ff878e7) with main (4b84fa9)1

Open in CodSpeed

Footnotes

  1. No successful run was found on main (174c1ac) during the generation of this report, so 4b84fa9 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

@greptile-apps

greptile-apps Bot commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR prevents streamed Responses API calls from being recorded with zero usage when a provider omits usage from a completed or incomplete terminal event.

  • Estimates input tokens from transformed Responses API messages.
  • Counts streamed text and function, custom-tool, and MCP argument deltas as output.
  • Preserves provider-reported usage when available.
  • Uses immutable copies when adding estimated usage and when upgrading dictionary responses.
  • Keeps serialization and estimation failures non-fatal.

Confidence Score: 5/5

The PR appears safe to merge; the previously reported billing, structured-input, tool-output, dictionary-response, and immutability issues are resolved.

No actionable new failure remains, and every previous Greptile thread is resolved with the current code implementing immutable response copies and preserving the corrected billing behavior.

Important Files Changed
Filename Overview
litellm/responses/streaming_iterator.py Adds best-effort streamed usage estimation, immutable terminal-response upgrades, cost stamping, and non-fatal cache serialization handling.
tests/test_litellm/responses/test_streaming_iterator.py Adds regression coverage for text, tool arguments, multimodal input, provider usage preservation, immutable copies, dictionary responses, and serialization failures.

Reviews (7): Last reviewed commit: "refactor(responses): copy the terminal e..." | Re-trigger Greptile

Comment thread litellm/responses/streaming_iterator.py Outdated
Comment thread litellm/responses/utils.py Outdated
@codecov

codecov Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.56098% with 1 line in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
litellm/responses/streaming_iterator.py 97.56% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

kerry-berri and others added 2 commits September 16, 2026 01:02
…reamed usage fallback

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ng raises

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread litellm/responses/streaming_iterator.py Outdated
kerry-berri and others added 2 commits September 16, 2026 01:24
…amed usage fallback

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ng litellm

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@kerry-berri

Copy link
Copy Markdown
Contributor

@greptile run

@kerry-berri

Copy link
Copy Markdown
Contributor

@BugBot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/responses/streaming_iterator.py Outdated
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@kerry-berri

Copy link
Copy Markdown
Contributor

@greptile run

@kerry-berri

Copy link
Copy Markdown
Contributor

@BugBot run

Comment thread litellm/responses/streaming_iterator.py Outdated

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/responses/streaming_iterator.py
…e is billed

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@kerry-berri

Copy link
Copy Markdown
Contributor

@greptile run

@kerry-berri

Copy link
Copy Markdown
Contributor

@BugBot run

Comment thread litellm/responses/streaming_iterator.py Outdated

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/responses/streaming_iterator.py Outdated
kerry-berri and others added 2 commits September 16, 2026 03:47
…d the cache dump

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…bed chunks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@kerry-berri

Copy link
Copy Markdown
Contributor

@greptile run

@kerry-berri

Copy link
Copy Markdown
Contributor

@BugBot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit ff878e7. Configure here.

@kerry-berri
kerry-berri merged commit 5960881 into main Sep 16, 2026
90 checks passed
@kerry-berri
kerry-berri deleted the litellm_fix_responses_stream_absent_usage_recount branch September 16, 2026 04:19
yuneng-berri added a commit that referenced this pull request Sep 16, 2026
…or tests

PR #41337 merged without the CircleCI lane and its immutable terminal-response rewrite reads response.usage and returns model_copy() results, which Mock(spec=...) stand-ins cannot satisfy: two tests raised AttributeError on usage and two compared a child Mock's type to the event enum. The four tests now construct real ResponsesAPIResponse and event objects
yuneng-berri added a commit that referenced this pull request Sep 17, 2026
leilei3167 pushed a commit to leilei3167/litellm that referenced this pull request Sep 17, 2026
…iAI#39996 and BerriAI#41310

Every one of these fails on main's own scheduled CircleCI run with the same
assertion as on any PR, and each traces to a merged behavior change that
never updated the test that pinned the old behavior

- tests/integration/_support/client.py: BerriAI#41311 made /key/info serve deleted
  keys from the archive with status deleted, so the scenario teardown asserts
  the live row is gone and the readback reports deleted instead of a 404.
  This alone accounts for nine integration-management and one
  integration-providers failure
- tests/integration/authorization/test_warmed_policy.py: BerriAI#39996 made team
  admins unable to edit any team field unless a proxy admin allow-lists it,
  and tpm_limit is the only field it accepts today. The demotion test now
  enables tpm_limit for the scenario and edits that instead of team_alias
- tests/llm_responses_api_testing/test_base_responses_api_streaming_iterator.py:
  BerriAI#41337 reads usage off the terminal response and copies the event when it
  is missing, which a Mock(spec=ResponsesAPIResponse) cannot survive. The
  four mocks now carry a usage object
- tests/test_openai_endpoints.py: BerriAI#41310 lengthened the access-denied
  message, and the test matched against the ExceptionInfo repr, which
  saferepr truncates in the middle. It now matches the exception text
- tests/local_testing/test_text_completion.py: Together no longer serves
  Qwen2-1.5B serverless, the cheapest cost-map row. The test mocks the
  completions call and asserts the request litellm builds, so a vendor
  catalog rotation cannot fail it again

test_router_fallbacks_with_cooldowns_and_dynamic_credentials is deliberately
untouched: it passes and fails on main with identical code, and the failing
path is a product question about whether dynamic-credential 429s cool down
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants