fix(responses): recount tokens when a streamed response completes without usage - #41337
Conversation
…t omits usage Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
🤖 Devin AI EngineerI'll be helping with this pull request! Here's what you should know: ✅ I will automatically:
Note: I can only respond to comments from users who have write access to this repository. ⚙️ Control Options:
|
|
|
Greptile SummaryThis PR prevents streamed Responses API calls from being recorded with zero usage when a provider omits usage from a completed or incomplete terminal event.
Confidence Score: 5/5The PR appears safe to merge; the previously reported billing, structured-input, tool-output, dictionary-response, and immutability issues are resolved. No actionable new failure remains, and every previous Greptile thread is resolved with the current code implementing immutable response copies and preserving the corrected billing behavior. Important Files Changed
Reviews (7): Last reviewed commit: "refactor(responses): copy the terminal e..." | Re-trigger Greptile |
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
…reamed usage fallback Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ng raises Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…amed usage fallback Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ng litellm Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
@greptile run |
|
@BugBot run |
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
@greptile run |
|
@BugBot run |
…e is billed Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
@greptile run |
|
@BugBot run |
…d the cache dump Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…bed chunks Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
@greptile run |
|
@BugBot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit ff878e7. Configure here.
…or tests PR #41337 merged without the CircleCI lane and its immutable terminal-response rewrite reads response.usage and returns model_copy() results, which Mock(spec=...) stand-ins cannot satisfy: two tests raised AttributeError on usage and two compared a child Mock's type to the event enum. The four tests now construct real ResponsesAPIResponse and event objects
…iAI#39996 and BerriAI#41310 Every one of these fails on main's own scheduled CircleCI run with the same assertion as on any PR, and each traces to a merged behavior change that never updated the test that pinned the old behavior - tests/integration/_support/client.py: BerriAI#41311 made /key/info serve deleted keys from the archive with status deleted, so the scenario teardown asserts the live row is gone and the readback reports deleted instead of a 404. This alone accounts for nine integration-management and one integration-providers failure - tests/integration/authorization/test_warmed_policy.py: BerriAI#39996 made team admins unable to edit any team field unless a proxy admin allow-lists it, and tpm_limit is the only field it accepts today. The demotion test now enables tpm_limit for the scenario and edits that instead of team_alias - tests/llm_responses_api_testing/test_base_responses_api_streaming_iterator.py: BerriAI#41337 reads usage off the terminal response and copies the event when it is missing, which a Mock(spec=ResponsesAPIResponse) cannot survive. The four mocks now carry a usage object - tests/test_openai_endpoints.py: BerriAI#41310 lengthened the access-denied message, and the test matched against the ExceptionInfo repr, which saferepr truncates in the middle. It now matches the exception text - tests/local_testing/test_text_completion.py: Together no longer serves Qwen2-1.5B serverless, the cheapest cost-map row. The test mocks the completions call and asserts the request litellm builds, so a vendor catalog rotation cannot fail it again test_router_fallbacks_with_cooldowns_and_dynamic_credentials is deliberately untouched: it passes and fails on main with identical code, and the failing path is a product question about whether dynamic-credential 429s cool down
TLDR
Problem this solves:
/v1/responsescall whose completed event has nousageis billed $0How it solves it:
response.completed/response.incompletewithusage: null, estimate usage from the request input and the streamed outputresponsemay be a typedResponsesAPIResponseor, when the OpenAI transformation falls back tomodel_construct, a plain dict; a dict is upgraded to a typed object and the terminal event is re-issued as a copy carrying it, so the estimate, cost stamping and logging all run on it instead of crashing or billing $0. Nothing is mutated in place, and the response cache write tolerates an upgraded object that cannot be serialized instead of failing the streamUser Flow
Before: a developer streaming from a Responses-compatible upstream that omits usage sees the call logged for free
"model": "gpt-5.6","stream": trueresponse.completedwith"usage": nullAfter: the same stream is billed from a token recount
"stream": trueresponse.completednow carries ausageobject with estimated input and output tokensRelevant issues
Found by the scripted-provider cost suite in #41328
Affected release
Linear ticket
Resolves LIT-7840
Pre-Submission checklist
uv run pytest tests/test_litellm/responses/test_streaming_iterator.py tests/test_litellm/responses/test_responses_utils.pyplus the wholetests/test_litellm/responsesdirectory (711 passed)Screenshots / Proof of Fix
Regression tests stream output deltas followed by a
response.completedevent withusage: nulland assert the completed response carries input and output token counts greater than zero withtotal = input + output; the same for a function-call-only stream; a multimodal input counts fewer tokens than its JSON serialization; custom-tool and MCP argument deltas are counted like function-call arguments; a completed event that does carry usage is left untouched; amodel_constructed completed event whoseresponseis a plain dict ends up typed with estimated usage and a stamped cost, with the cost calculator called on it and the yielded terminal event being the re-issued copy; a typed response that already has usage is passed through as the same object while one without usage gets a copy and the original is left untouched; an upgraded response whose JSON dump raises is skipped by the cache write without failing the stream; and a malformed input that makes the estimate raise leaves usage unset without failing the stream, with no monkeypatching involved. Live before/after against the scripted upstream will follow once #41328 lands, where the responsesstream_no_usagecase flips fromexpect_zero_billto a non-zero billType
🐛 Bug Fix
Caveats (if any)
Medium
response.failedis left alone and stays $0Low
Note
Medium Risk
Changes proxy-side token estimation and spend calculation for Responses streaming when providers omit usage; incorrect estimates could misbill, though provider usage still wins and estimate failures fall back to $0 without failing the stream.
Overview
Fixes $0 billing when a streamed
/v1/responsescall finishes withusage: nullonresponse.completedorresponse.incomplete.The streaming iterator now accumulates output text plus function-call, custom-tool, and MCP argument deltas, then on terminal events estimates
ResponseAPIUsageviatoken_counter(request input converted to chat messages first, so multimodal input is not counted as raw JSON). Provider-reported usage is unchanged;response.failedis not estimated. Estimates are attached through_billed_terminal_response, which copies the terminal event when needed so cost stamping and logging see typed usage without mutating the original object; plain-dict responses frommodel_constructare upgraded the same way.Cache persistence uses
_dump_json_safelyso an unserializable completed response skips the cache write instead of breaking the stream. Regression tests cover text-only and tool-only streams, multimodal input, estimate failures, dict responses, and cache serialization edge cases.Reviewed by Cursor Bugbot for commit ff878e7. Bugbot is set up for automated code reviews on this repo. Configure here.
Link to Devin session: https://app.devin.ai/sessions/25b6f01a2b594d44a5e2488dc54d010c
Open in Devin Desktop: https://app.devin.ai/desktop/session/25b6f01a2b594d44a5e2488dc54d010c?variant=devin