fix(streaming): report response_cost and Anthropic citations from stream_chunk_builder - #38696
Conversation
Greptile SummaryThe PR restores response-cost metadata on assembled streaming responses and preserves streamed citations.
Confidence Score: 5/5The PR appears safe to merge. No blocking failure remains; the previously reported Databricks citation nesting defect is corrected at the current head.
|
| Filename | Overview |
|---|---|
| litellm/main.py | Adds response-cost calculation and citation aggregation; the Databricks nesting issue from the previous review is fixed. |
| litellm/litellm_core_utils/streaming_handler.py | Passes the logging object into partial-stream assembly so failure spend uses logging-aware pricing. |
| tests/test_litellm/litellm_core_utils/test_streaming_handler.py | Extends partial-stream recovery coverage for corrected-model pricing. |
| tests/test_litellm/test_main.py | Covers known, unknown, aliased, and logging-aware stream cost behavior. |
| tests/test_litellm/test_stream_chunk_builder_citations.py | Verifies Anthropic citation collection and corrected Databricks block-list nesting. |
Reviews (2): Last reviewed commit: "fix(streaming): join block-list citation..." | Re-trigger Greptile
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
dlowzzxx
left a comment
There was a problem hiding this comment.
Substantive Technical Code Review: BerriAI/litellm #38696
Target Repository: BerriAI/litellm
PR Number: #38696
PR Title: fix(streaming): report response_cost and Anthropic citations from stream_chunk_builder
Author: krrishdholakia
Target Branch: main
Reviewer: Teamwork Ecosystem PR Reviewer
Status: REQUEST_CHANGES
1. Executive Summary & PR Metadata
This pull request addresses two long-standing regressions in LiteLLM's stream reconstruction pipeline (stream_chunk_builder):
- Loss of Streamed Anthropic Citations: When streaming responses from Anthropic Claude (which yields citations chunk-by-chunk in
delta.provider_specific_fields.citation),stream_chunk_builderpreviously merged dictionary fields with a simple key-assignment loop. This caused subsequent citation chunks to overwrite prior chunks, discarding all but the final citation in the mergedModelResponse. - Missing
response_coston Streaming Completion: In streaming mode,response._hidden_params["response_cost"]remained unset unless an external logging object calculated it. The PR introduces fallback calculation usinglitellm.completion_cost().
Cross-Audited PRs in Repository:
- PR #38690 (
fix(router): scrub fallback stamp keys in place and strip them at the proxy boundary): Fixes proxy metadata bucket detachment by scrubbing reserved router fields in place. - PR #38686 (
fix(proxy): reset a key's budget-window counters on spend reset): Fixes mid-window spend reset desynchronization across Redis, Prisma DB, and memory caches.
2. Architectural & Design Assessment
2.1 Adherence to LiteLLM Core Patterns
stream_chunk_builder in litellm/main.py is the central aggregator responsible for converting arbitrary sequences of provider streaming chunks into a single unified ModelResponse.
+-----------------------------------+
| Heterogeneous Stream Chunks |
| (OpenAI, Anthropic, Bedrock...) |
+-----------------+-----------------+
|
v
+-----------------+-----------------+
| stream_chunk_builder() |
| - Delta content concatenation |
| - Usage token accumulation |
| - Provider-specific field merge |
| - Cost calculation & hiding |
+-----------------+-----------------+
|
v
+-----------------+-----------------+
| Unified ModelResponse |
| - choices[0].message.content |
| - provider_specific_fields |
| - _hidden_params["response_cost"]|
+-----------------------------------+
2.2 Design Evaluation in PR #38696
- Citations Array Normalization: Aggregates all
delta.provider_specific_fields["citation"]occurrences into a nested list{"citations": [list(streamed_citations)]}, matching the schema produced by non-streaming Anthropic responses. - Hidden Param Decorator: Isolates cost calculation into
_stream_builder_response_cost()and_set_stream_builder_response_cost(), respecting existinglogging_objownership when available.
3. Correctness, Concurrency & Edge-Case Analysis
Finding 1 [CRITICAL]: Polymorphic Chunk Typing Failure (ModelResponseStream vs dict)
- Location:
litellm/main.py:8814-8840 - Mechanism:
In LiteLLM,stream_chunk_builderacceptschunks: list, which can contain standard Python dictionaries OR Pydantic/OpenAIModelResponseStreaminstances (as produced bylitellm.acompletion(..., stream=True)when iterated natively).
The PR's filter checks:provider_specific_chunks = [ chunk for chunk in chunks if ( isinstance(chunk, dict) and chunk.get("choices") and isinstance(chunk["choices"], list) and len(chunk["choices"]) > 0 and isinstance(chunk["choices"][0], dict) and chunk["choices"][0].get("delta") and isinstance(chunk["choices"][0]["delta"], dict) and "provider_specific_fields" in chunk["choices"][0]["delta"] ) ]
- Failure Mode:
IfchunkscontainsModelResponseStreaminstances,isinstance(chunk, dict)evaluates toFalsefor every element. As a result,provider_specific_chunksis empty, and ALL citations are silently dropped whenstream_chunk_builderis called on real Pydantic chunk streams! - Severity: Critical. Silently invalidates the feature when interacting with LiteLLM's standard response objects.
Finding 2 [MAJOR]: Hardcoded chunk["choices"][0] Multi-Choice (n > 1) and Empty Choices Hazard
-
Location:
litellm/main.py:8842-8846 -
Mechanism:
provider_field_dicts: Final = tuple( fields for chunk in provider_specific_chunks for fields in (chunk["choices"][0]["delta"]["provider_specific_fields"],) if isinstance(fields, dict) )
-
Failure Modes:
-
Multi-Choice Loss: For requests with
$n > 1$ (multiple completions generated concurrently), citations and provider fields on choices[1],[2], etc., are completely ignored. -
Index / Type Error: If a chunk has an empty
choiceslist (e.g. SSE keep-alive or comment chunks emitted by some gateways), direct indexing[0]can raiseIndexErrorunless filtered.
-
Multi-Choice Loss: For requests with
Finding 3 [MEDIUM]: Stale Provider Hint in _stream_builder_response_cost
- Location:
litellm/main.py:8590-8605 - Mechanism:
provider_hint: Final = response._hidden_params.get("custom_llm_provider") try: return litellm.completion_cost(completion_response=response, custom_llm_provider=provider_hint) except Exception: return None
- Failure Mode:
Whenstream_chunk_builderis called on chunks from models with provider prefixes (e.g.bedrock/anthropic.claude-3-5-sonnetorazure/gpt-4o),response._hidden_paramsdoes not contain"custom_llm_provider"unless explicitly injected. Whilecompletion_costattempts to split prefixes, ifresponse.modelwas normalized to the bare model name without the prefix, cost calculation fails and returnsNone.
4. Concrete Code Diff Recommendations
Apply the following patch to support both dict and ModelResponseStream chunk types, handle all choices safely, and ensure robust cost calculation:
--- a/litellm/main.py
+++ b/litellm/main.py
@@ -8813,32 +8813,44 @@ def stream_chunk_builder(
provider_specific_chunks = [
chunk
for chunk in chunks
- if (
- isinstance(chunk, dict)
- and chunk.get("choices")
- and isinstance(chunk["choices"], list)
- and len(chunk["choices"]) > 0
- and isinstance(chunk["choices"][0], dict)
- and chunk["choices"][0].get("delta")
- and isinstance(chunk["choices"][0]["delta"], dict)
- and "provider_specific_fields" in chunk["choices"][0]["delta"]
- )
+ if chunk is not None
]
if len(provider_specific_chunks) > 0:
- provider_field_dicts: Final = tuple(
- fields
- for chunk in provider_specific_chunks
- for fields in (chunk["choices"][0]["delta"]["provider_specific_fields"],)
- if isinstance(fields, dict)
- )
+ collected_fields: list[dict[str, object]] = []
+ for chunk in provider_specific_chunks:
+ # Extract choices whether chunk is a dict or a ModelResponseStream
+ choices = chunk.get("choices", []) if isinstance(chunk, dict) else getattr(chunk, "choices", [])
+ if not choices:
+ continue
+ for choice in choices:
+ delta = choice.get("delta") if isinstance(choice, dict) else getattr(choice, "delta", None)
+ if delta is None:
+ continue
+ fields = (
+ delta.get("provider_specific_fields")
+ if isinstance(delta, dict)
+ else getattr(delta, "provider_specific_fields", None)
+ )
+ if isinstance(fields, dict):
+ collected_fields.append(fields)
+
+ provider_field_dicts: Final = tuple(collected_fields)
streamed_citations: Final = tuple(
fields["citation"] for fields in provider_field_dicts if fields.get("citation") is not None
)
citation_fields: Final = (
{"citations": [list(streamed_citations)]} if streamed_citations else {} # mutable-ok: JSON dict field
)
combined_provider_fields: Final = { # mutable-ok: Message.provider_specific_fields is a plain dict field
key: value
for fields in (citation_fields, *provider_field_dicts)
for key, value in fields.items()
if key != "citation"
}5. Verification & Test Suite Recommendations
Missing Test Cases to Add:
-
test_stream_chunk_builder_citations_with_model_response_stream_objects: Test withModelResponseStreaminstances to verify Pydantic object traversal. -
test_stream_chunk_builder_citations_multi_choice: Test with$n=2$ streamed completions. -
test_stream_chunk_builder_empty_delta_keepalive: Test with SSE keep-alive chunks wherechoices: [].
6. Final Review Verdict
Verdict: REQUEST_CHANGES
Rationale: While the citation aggregation logic is sound for dictionary-based chunk mocks, it completely fails in production when supplied with native ModelResponseStream objects due to strict isinstance(chunk, dict) checks. Implementing the polymorphic extractor diff above resolves all failure modes.
|
@dlowzzxx This seems like an auto-generated review. All of that is pre-existing and not modified by this PR |
|
bugbot run |
| if isinstance(usage_cost, (int, float)): | ||
| return float(usage_cost) | ||
| if logging_obj is not None: | ||
| return None |
There was a problem hiding this comment.
Stamped stream cost skips corrected model
Medium Severity
_stream_builder_response_cost still copies usage.cost into _hidden_params["response_cost"] when a logging_obj is present. Mid-stream failure recovery can set usage.cost via _response_cost_calculator before it overwrites partial_response.model, then _response_cost_calculator trusts that stamped value and never re-prices with the corrected model, custom pricing, or later logging-pipeline adjustments.
Additional Locations (2)
Reviewed by Cursor Bugbot for commit 9b1b8e7. Configure here.
There was a problem hiding this comment.
Provider-reported usage.cost wins by design: response_cost_calculator returns it before custom pricing on the success path, so failure recovery pricing matches a completed stream
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
1 issue from previous review remains unresolved.
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 9b1b8e7. Configure here.
|
Bugbot Autofix prepared a fix for the issue found in the latest run.
Or push these changes by commenting: Preview (41b9a41667)diff --git a/litellm/main.py b/litellm/main.py
--- a/litellm/main.py
+++ b/litellm/main.py
@@ -8591,11 +8591,11 @@
def _stream_builder_response_cost(response: ModelResponse, logging_obj: Optional["Logging"]) -> float | None:
+ if logging_obj is not None:
+ return None
usage_cost: Final = getattr(getattr(response, "usage", None), "cost", None)
if isinstance(usage_cost, (int, float)):
return float(usage_cost)
- if logging_obj is not None:
- return None
provider_hint: Final = response._hidden_params.get( # pyright: ignore[reportPrivateUsage] # no public accessor
"custom_llm_provider"
)
diff --git a/tests/test_litellm/litellm_core_utils/test_streaming_handler.py b/tests/test_litellm/litellm_core_utils/test_streaming_handler.py
--- a/tests/test_litellm/litellm_core_utils/test_streaming_handler.py
+++ b/tests/test_litellm/litellm_core_utils/test_streaming_handler.py
@@ -3484,6 +3484,52 @@
assert logging_obj.model_call_details["response_cost"] == pytest.approx(expected)
+def test_stream_chunk_builder_defers_response_cost_when_logging_obj_present(monkeypatch):
+ """Regression: passing logging_obj into stream_chunk_builder must leave the
+ downstream logging pipeline in charge of pricing. Stamping response_cost
+ on _hidden_params from the raw chunk model would short-circuit the
+ logging-object recalculator and bake the wrong price into failure spend.
+ """
+ monkeypatch.setattr(litellm, "include_cost_in_streaming_usage", True)
+ logging_obj = Logging(
+ model="gpt-4o-mini",
+ messages=[{"role": "user", "content": "hi"}],
+ stream=True,
+ call_type="completion",
+ start_time=time.time(),
+ litellm_call_id="stamped-cost-defer",
+ function_id="1245",
+ )
+ logging_obj.model_call_details["custom_llm_provider"] = "openai"
+ logging_obj.optional_params = {}
+
+ chunks = [
+ ModelResponseStream(
+ id="chatcmpl-defer-1",
+ created=1742056047,
+ model="claude-opus-4-5",
+ object="chat.completion.chunk",
+ choices=[
+ StreamingChoices(
+ finish_reason=None,
+ index=0,
+ delta=Delta(content="hi", role="assistant"),
+ )
+ ],
+ usage=Usage(prompt_tokens=40, completion_tokens=5, total_tokens=45),
+ )
+ ]
+
+ partial = litellm.stream_chunk_builder(
+ chunks=chunks,
+ messages=[{"role": "user", "content": "hi"}],
+ logging_obj=logging_obj,
+ )
+
+ assert partial is not None
+ assert "response_cost" not in partial._hidden_params
+
+
def test_record_partial_usage_for_failure_carries_up_openai_style_cached_tokens():
recovered = Usage(
prompt_tokens=1000,You can send follow-ups to the cloud agent here. |
4ebedf9
into
litellm_internal_staging



TLDR
Problem this solves:
track_usagerecordscost: Nonefor every streamed callHow it solves it:
citationslistUser Flow
Before: a DSPy developer who turns on streaming loses cost tracking and citations
dspy.streamify(dspy.Predict("question -> answer"))withtrack_usage=Trueanddspy.LM("anthropic/claude-opus-5", cache=False)"stream": trueand the answer streams back finelm.history[-1]["cost"]readsNone, so their spend tracking records nothing, while the identical non-streamed call records 0.00306After: the same streamed program reports real spend and keeps its citations
dspy.streamify(dspy.Predict("question -> answer"))withtrack_usage=Trueanddspy.LM("anthropic/claude-opus-5", cache=False)"stream": trueand the answer streams back finelm.history[-1]["cost"]now reads 0.00318, in line with the non-streamed callRelevant issues
Linear ticket
Resolves LIT-6376
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Screenshots / Proof of Fix
Legs: before at the merge base 936e07b, after at this PR's head 9b1b8e7. Same client both sides: DSPy 3.3.1 driving real Anthropic, OpenAI, Gemini, and Bedrock APIs, no mocks. SDK cases call providers directly. Proxy cases boot one proxy instance from the leg's commit with
--num_workers 2(no DB) on a random localhost port, with the client litellm at the same commit and models exposed as OpenAI-compatible aliases (gpt-5.6,claude-opus-5; the after leg also drivesgemini-3.7-flash).stream_chunk_builderonly joins chat-completions-shaped streams (SDK joining and proxy chat logging); /v1/messages and /v1/responses have their own assemblers and never ride this code, so those endpoints are out of scope hereStreamed cost, run with DSPy exactly as an end user would:
Same thing through a LiteLLM proxy, pointing DSPy at the alias:
Streamed citations, run against the live Anthropic API:
Before (936e07b)
Streamed cost lost
anthropic/claude-opus-5,gpt-5.6,gemini/gemini-3.7-flash,bedrock/us.anthropic.claude-sonnet-5)cost: Noneon all four, while the identical non-streamed calls priced finecost: Noneon both thegpt-5.6andclaude-opus-5aliases, while the identical non-streamed calls read 0.001164 and 0.003125Streamed citations lost
citations: Noneandresponse_cost: Noneon the joined responseAfter (9b1b8e7)
Streamed cost lost
gpt-5.60.001148 (nonstream 0.001164),claude-opus-50.00295 (nonstream 0.003125),gemini-3.7-flash0.000996 (nonstream 0.001065)Streamed citations lost
char_locationon the document) instead ofNoneDatabricks streamed citations A/B (SDK direct)
Same joined-stream flow against the real Databricks workspace, model
databricks/databricks-claude-sonnet-5, no mocks:null; the citations never reach the joined message[[c1, c2]]withchar_locationentries, no double nesting. This live run emitted singularcitationdeltas; the block-list delta shape from the original report stays locked by the regression tests intests/test_litellm/test_stream_chunk_builder_citations.pyObservations from the runs:
openai/aliases rejectreasoning_effortclient-side; pre-existing, left alonesupported_text; pre-existing, left aloneLive PR risk: CHECKED at 9b1b8e7. Two findings were fixed in-PR, each with a regression test: mid-stream failure spend priced off the raw chunk model, and Databricks block-list citation deltas double-nested. CircleCI ran under the run-ci label; the logging_testing red is pre-existing (it fails on staging commits c11a1f0 and 3daf7a3 and on unrelated PR #38440, with identical failure sets at head vs merge base locally). A follow-up live Databricks A/B (above) closed the citation-join gap that had only unit-level coverage
Type
🐛 Bug Fix
Caveats (if any)
Low
supported_text; non-streamed ones carry itFinal Attestation
Note
Medium Risk
Changes billing-adjacent assembly (response_cost, partial-failure spend) and provider metadata merging; behavior is covered by new tests but affects all consumers of joined streams.
Overview
Joined streamed chat responses now expose spend and citations that were previously missing for tools like DSPy
track_usageand Anthropic document citations.stream_chunk_builderstampsresponse_coston_hidden_paramswhen it can derive a price: providerusage.costwins; otherwise it prices viacompletion_cost/ the model map (with a proxy-alias fallback). When alogging_objis passed and usage has no cost, hiddenresponse_costis left unset so the existing logging cost pipeline applies (e.g.include_cost_in_streaming_usage).Citation deltas in
provider_specific_fieldsare aggregated into a finalcitationslist (singularcitationkeys removed), including block-list citation shapes without extra nesting.Mid-stream failure spend recovery passes
logging_objintostream_chunk_builderso partial failures use the corrected model and full cost calculator instead of raw chunk model pricing.Reviewed by Cursor Bugbot for commit 9b1b8e7. Bugbot is set up for automated code reviews on this repo. Configure here.