Skip to content

fix(streaming): report response_cost and Anthropic citations from stream_chunk_builder - #38696

Merged
mateo-berri merged 4 commits into
litellm_internal_stagingfrom
litellm_lit6376_dspy_streaming
Aug 28, 2026
Merged

mateo-berri merged 4 commits into
litellm_internal_stagingfrom
litellm_lit6376_dspy_streaming

Conversation

@mateo-berri

@mateo-berri mateo-berri commented Aug 28, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Responses joined from streamed chunks carry no response cost
  • DSPy's track_usage records cost: None for every streamed call
  • Streamed Anthropic citations are dropped when chunks are joined

How it solves it:

  • The joined response now carries its response cost in hidden params
  • Provider-reported usage cost wins; otherwise litellm prices the usage, falling back to the bare model name in the model map when the declared provider can't price it (models aliased through a LiteLLM proxy)
  • Citation deltas are collected into the final message's citations list
  • Block-list citation deltas (Databricks) join without extra nesting, matching the non-streamed shape
  • Internal joins that feed the proxy's logging pipeline (mid-stream failure spend recovery) now declare their logging object, so failure spend keeps coming from the full cost pipeline (custom pricing, corrected model) instead of a cost stamped from the raw chunk model

User Flow

Before: a DSPy developer who turns on streaming loses cost tracking and citations

  1. They wrap their program in dspy.streamify(dspy.Predict("question -> answer")) with track_usage=True and dspy.LM("anthropic/claude-opus-5", cache=False)
  2. DSPy sends POST https://api.anthropic.com/v1/messages with "stream": true and the answer streams back fine
  3. lm.history[-1]["cost"] reads None, so their spend tracking records nothing, while the identical non-streamed call records 0.00306
  4. With a document attached and citations enabled, they watch citation deltas arrive in the stream, yet the final prediction's citations come back empty, while the identical non-streamed call returns the citation

After: the same streamed program reports real spend and keeps its citations

  1. They wrap their program in dspy.streamify(dspy.Predict("question -> answer")) with track_usage=True and dspy.LM("anthropic/claude-opus-5", cache=False)
  2. DSPy sends POST https://api.anthropic.com/v1/messages with "stream": true and the answer streams back fine
  3. lm.history[-1]["cost"] now reads 0.00318, in line with the non-streamed call
  4. The final prediction carries every citation that streamed by

Relevant issues

Linear ticket

Resolves LIT-6376

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Screenshots / Proof of Fix

Legs: before at the merge base 936e07b, after at this PR's head 9b1b8e7. Same client both sides: DSPy 3.3.1 driving real Anthropic, OpenAI, Gemini, and Bedrock APIs, no mocks. SDK cases call providers directly. Proxy cases boot one proxy instance from the leg's commit with --num_workers 2 (no DB) on a random localhost port, with the client litellm at the same commit and models exposed as OpenAI-compatible aliases (gpt-5.6, claude-opus-5; the after leg also drives gemini-3.7-flash). stream_chunk_builder only joins chat-completions-shaped streams (SDK joining and proxy chat logging); /v1/messages and /v1/responses have their own assemblers and never ride this code, so those endpoints are out of scope here

Streamed cost, run with DSPy exactly as an end user would:

python - <<'PY'
import asyncio
import dspy

lm = dspy.LM("anthropic/claude-opus-5", cache=False)
dspy.configure(lm=lm, track_usage=True)
program = dspy.streamify(dspy.Predict("question -> answer"))

async def run():
    async for _ in program(question="Why is the sky blue? One sentence."):
        pass

asyncio.run(run())
print("cost:", lm.history[-1]["cost"])
PY

Same thing through a LiteLLM proxy, pointing DSPy at the alias:

lm = dspy.LM("openai/claude-opus-5", api_base="http://localhost:<port>/v1", api_key=MASTER_KEY, cache=False)

Streamed citations, run against the live Anthropic API:

python - <<'PY'
import litellm

doc = {
    "type": "document",
    "source": {"type": "text", "media_type": "text/plain",
               "data": "The Eiffel Tower was completed in 1889. It is 330 metres tall."},
    "title": "facts",
    "citations": {"enabled": True},
}
chunks = list(litellm.completion(
    model="anthropic/claude-opus-5",
    messages=[{"role": "user", "content": [doc, {"type": "text", "text": "How tall is the Eiffel Tower and when was it completed? Cite the document."}]}],
    stream=True,
))
built = litellm.stream_chunk_builder(chunks)
print("citations:", (built.choices[0].message.provider_specific_fields or {}).get("citations"))
print("response_cost:", built._hidden_params.get("response_cost"))
PY

Before (936e07b)

Streamed cost lost

  1. Ran the DSPy snippet per provider (anthropic/claude-opus-5, gpt-5.6, gemini/gemini-3.7-flash, bedrock/us.anthropic.claude-sonnet-5)
  2. Observed cost: None on all four, while the identical non-streamed calls priced fine
  3. Through the proxy (2 workers), streamed cost: None on both the gpt-5.6 and claude-opus-5 aliases, while the identical non-streamed calls read 0.001164 and 0.003125

Streamed citations lost

  1. Ran the citations snippet: 3 citation deltas arrived in the stream, then citations: None and response_cost: None on the joined response
  2. DSPy's streaming Citations program direct to Anthropic ended with 0 citations on the final prediction, where the identical non-streamed run returned 1
  3. Through the proxy, the joined response likewise carried 0 citations

After (9b1b8e7)

Streamed cost lost

  1. Ran the same DSPy snippet per provider
  2. Observed real spend on all four: anthropic 0.00355, openai 0.001204, gemini 0.00108375, bedrock 0.001276
  3. Through the proxy (2 workers), streamed cost now reads on the aliases: gpt-5.6 0.001148 (nonstream 0.001164), claude-opus-5 0.00295 (nonstream 0.003125), gemini-3.7-flash 0.000996 (nonstream 0.001065)

Streamed citations lost

  1. Ran the same citations snippet: 1 citation delta arrived and the joined response carries 1 citation (char_location on the document) instead of None
  2. DSPy's streaming Citations program direct to Anthropic ended with 1 citation on the final prediction, matching the non-streamed run (cost 0.007655 streamed, 0.00763 non-streamed)
  3. Through the proxy, the joined response carries 1 citation (non-streamed also 1) and the streamed DSPy Citations run reports cost 0.009825

Databricks streamed citations A/B (SDK direct)

Same joined-stream flow against the real Databricks workspace, model databricks/databricks-claude-sonnet-5, no mocks:

python - <<'PY'
import json
import litellm

document = {
    "type": "document",
    "source": {
        "type": "text",
        "media_type": "text/plain",
        "data": "The sky is blue because of Rayleigh scattering. Grass is green because of chlorophyll.",
    },
    "title": "Science facts",
    "citations": {"enabled": True},
}

chunks = list(
    litellm.completion(
        model="databricks/databricks-claude-sonnet-5",
        messages=[{"role": "user", "content": [document, {"type": "text", "text": "Why is the sky blue, and why is grass green? Cite the document for each fact."}]}],
        stream=True,
    )
)
built = litellm.stream_chunk_builder(chunks)
print(json.dumps((built.choices[0].message.provider_specific_fields or {}).get("citations"), default=str))
PY
  1. Before (936e07b): the stream delivered 2 citation deltas, and the joined response printed null; the citations never reach the joined message
  2. After (9b1b8e7): both citations survive as one per-block list, [[c1, c2]] with char_location entries, no double nesting. This live run emitted singular citation deltas; the block-list delta shape from the original report stays locked by the regression tests in tests/test_litellm/test_stream_chunk_builder_citations.py

Observations from the runs:

  • Proxy openai/ aliases reject reasoning_effort client-side; pre-existing, left alone
  • Streamed citation dicts lack supported_text; pre-existing, left alone
  • Streaming responses carry no plain response-cost header; pre-existing, left alone
  • DSPy's own citation collection masked the loss via proxy; unaffected
  • Alias missing from the model map streams cost 0.0
  • litellm's wrapper logging already stamps 0.0 there; pre-existing, unchanged

Live PR risk: CHECKED at 9b1b8e7. Two findings were fixed in-PR, each with a regression test: mid-stream failure spend priced off the raw chunk model, and Databricks block-list citation deltas double-nested. CircleCI ran under the run-ci label; the logging_testing red is pre-existing (it fails on staging commits c11a1f0 and 3daf7a3 and on unrelated PR #38440, with identical failure sets at head vs merge base locally). A follow-up live Databricks A/B (above) closed the citation-join gap that had only unit-level coverage

Type

🐛 Bug Fix

Caveats (if any)

Low

  • Streamed citation dicts lack supported_text; non-streamed ones carry it
  • Without a provider-reported usage cost, pricing falls back to litellm's model map, and a model aliased through a proxy prices by its bare model name, which assumes the alias matches the upstream model
  • An alias matching nothing in the model map prices streamed calls at 0.0 (litellm's default zero-cost info for unknown OpenAI-compatible models; the non-streamed call reads the real cost from the proxy's cost header, which streams don't carry)

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Note

Medium Risk
Changes billing-adjacent assembly (response_cost, partial-failure spend) and provider metadata merging; behavior is covered by new tests but affects all consumers of joined streams.

Overview
Joined streamed chat responses now expose spend and citations that were previously missing for tools like DSPy track_usage and Anthropic document citations.

stream_chunk_builder stamps response_cost on _hidden_params when it can derive a price: provider usage.cost wins; otherwise it prices via completion_cost / the model map (with a proxy-alias fallback). When a logging_obj is passed and usage has no cost, hidden response_cost is left unset so the existing logging cost pipeline applies (e.g. include_cost_in_streaming_usage).

Citation deltas in provider_specific_fields are aggregated into a final citations list (singular citation keys removed), including block-list citation shapes without extra nesting.

Mid-stream failure spend recovery passes logging_obj into stream_chunk_builder so partial failures use the corrected model and full cost calculator instead of raw chunk model pricing.

Reviewed by Cursor Bugbot for commit 9b1b8e7. Bugbot is set up for automated code reviews on this repo. Configure here.

@greptile-apps

greptile-apps Bot commented Aug 28, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

The PR restores response-cost metadata on assembled streaming responses and preserves streamed citations.

  • Prices joined stream usage while allowing logging-aware internal joins to defer to the full cost pipeline.
  • Aggregates Anthropic citation deltas and preserves Databricks citation block shape.
  • Adds regression coverage for stream cost calculation, partial-failure spend recovery, and citation assembly.

Confidence Score: 5/5

The PR appears safe to merge.

No blocking failure remains; the previously reported Databricks citation nesting defect is corrected at the current head.

Important Files Changed

Filename Overview
litellm/main.py Adds response-cost calculation and citation aggregation; the Databricks nesting issue from the previous review is fixed.
litellm/litellm_core_utils/streaming_handler.py Passes the logging object into partial-stream assembly so failure spend uses logging-aware pricing.
tests/test_litellm/litellm_core_utils/test_streaming_handler.py Extends partial-stream recovery coverage for corrected-model pricing.
tests/test_litellm/test_main.py Covers known, unknown, aliased, and logging-aware stream cost behavior.
tests/test_litellm/test_stream_chunk_builder_citations.py Verifies Anthropic citation collection and corrected Databricks block-list nesting.

Reviews (2): Last reviewed commit: "fix(streaming): join block-list citation..." | Re-trigger Greptile

Comment thread litellm/main.py
@codecov

codecov Bot commented Aug 28, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@dlowzzxx dlowzzxx left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Substantive Technical Code Review: BerriAI/litellm #38696

Target Repository: BerriAI/litellm
PR Number: #38696
PR Title: fix(streaming): report response_cost and Anthropic citations from stream_chunk_builder
Author: krrishdholakia
Target Branch: main
Reviewer: Teamwork Ecosystem PR Reviewer
Status: REQUEST_CHANGES


1. Executive Summary & PR Metadata

This pull request addresses two long-standing regressions in LiteLLM's stream reconstruction pipeline (stream_chunk_builder):

  1. Loss of Streamed Anthropic Citations: When streaming responses from Anthropic Claude (which yields citations chunk-by-chunk in delta.provider_specific_fields.citation), stream_chunk_builder previously merged dictionary fields with a simple key-assignment loop. This caused subsequent citation chunks to overwrite prior chunks, discarding all but the final citation in the merged ModelResponse.
  2. Missing response_cost on Streaming Completion: In streaming mode, response._hidden_params["response_cost"] remained unset unless an external logging object calculated it. The PR introduces fallback calculation using litellm.completion_cost().

Cross-Audited PRs in Repository:

  • PR #38690 (fix(router): scrub fallback stamp keys in place and strip them at the proxy boundary): Fixes proxy metadata bucket detachment by scrubbing reserved router fields in place.
  • PR #38686 (fix(proxy): reset a key's budget-window counters on spend reset): Fixes mid-window spend reset desynchronization across Redis, Prisma DB, and memory caches.

2. Architectural & Design Assessment

2.1 Adherence to LiteLLM Core Patterns

stream_chunk_builder in litellm/main.py is the central aggregator responsible for converting arbitrary sequences of provider streaming chunks into a single unified ModelResponse.

                  +-----------------------------------+
                  |  Heterogeneous Stream Chunks      |
                  |  (OpenAI, Anthropic, Bedrock...)  |
                  +-----------------+-----------------+
                                    |
                                    v
                  +-----------------+-----------------+
                  |      stream_chunk_builder()       |
                  |  - Delta content concatenation    |
                  |  - Usage token accumulation       |
                  |  - Provider-specific field merge  |
                  |  - Cost calculation & hiding      |
                  +-----------------+-----------------+
                                    |
                                    v
                  +-----------------+-----------------+
                  |       Unified ModelResponse       |
                  |  - choices[0].message.content     |
                  |  - provider_specific_fields       |
                  |  - _hidden_params["response_cost"]|
                  +-----------------------------------+

2.2 Design Evaluation in PR #38696

  • Citations Array Normalization: Aggregates all delta.provider_specific_fields["citation"] occurrences into a nested list {"citations": [list(streamed_citations)]}, matching the schema produced by non-streaming Anthropic responses.
  • Hidden Param Decorator: Isolates cost calculation into _stream_builder_response_cost() and _set_stream_builder_response_cost(), respecting existing logging_obj ownership when available.

3. Correctness, Concurrency & Edge-Case Analysis

Finding 1 [CRITICAL]: Polymorphic Chunk Typing Failure (ModelResponseStream vs dict)

  • Location: litellm/main.py:8814-8840
  • Mechanism:
    In LiteLLM, stream_chunk_builder accepts chunks: list, which can contain standard Python dictionaries OR Pydantic/OpenAI ModelResponseStream instances (as produced by litellm.acompletion(..., stream=True) when iterated natively).
    The PR's filter checks:
    provider_specific_chunks = [
        chunk for chunk in chunks
        if (
            isinstance(chunk, dict)
            and chunk.get("choices")
            and isinstance(chunk["choices"], list)
            and len(chunk["choices"]) > 0
            and isinstance(chunk["choices"][0], dict)
            and chunk["choices"][0].get("delta")
            and isinstance(chunk["choices"][0]["delta"], dict)
            and "provider_specific_fields" in chunk["choices"][0]["delta"]
        )
    ]
  • Failure Mode:
    If chunks contains ModelResponseStream instances, isinstance(chunk, dict) evaluates to False for every element. As a result, provider_specific_chunks is empty, and ALL citations are silently dropped when stream_chunk_builder is called on real Pydantic chunk streams!
  • Severity: Critical. Silently invalidates the feature when interacting with LiteLLM's standard response objects.

Finding 2 [MAJOR]: Hardcoded chunk["choices"][0] Multi-Choice (n > 1) and Empty Choices Hazard

  • Location: litellm/main.py:8842-8846
  • Mechanism:
    provider_field_dicts: Final = tuple(
        fields
        for chunk in provider_specific_chunks
        for fields in (chunk["choices"][0]["delta"]["provider_specific_fields"],)
        if isinstance(fields, dict)
    )
  • Failure Modes:
    1. Multi-Choice Loss: For requests with $n &gt; 1$ (multiple completions generated concurrently), citations and provider fields on choices [1], [2], etc., are completely ignored.
    2. Index / Type Error: If a chunk has an empty choices list (e.g. SSE keep-alive or comment chunks emitted by some gateways), direct indexing [0] can raise IndexError unless filtered.

Finding 3 [MEDIUM]: Stale Provider Hint in _stream_builder_response_cost

  • Location: litellm/main.py:8590-8605
  • Mechanism:
    provider_hint: Final = response._hidden_params.get("custom_llm_provider")
    try:
        return litellm.completion_cost(completion_response=response, custom_llm_provider=provider_hint)
    except Exception:
        return None
  • Failure Mode:
    When stream_chunk_builder is called on chunks from models with provider prefixes (e.g. bedrock/anthropic.claude-3-5-sonnet or azure/gpt-4o), response._hidden_params does not contain "custom_llm_provider" unless explicitly injected. While completion_cost attempts to split prefixes, if response.model was normalized to the bare model name without the prefix, cost calculation fails and returns None.

4. Concrete Code Diff Recommendations

Apply the following patch to support both dict and ModelResponseStream chunk types, handle all choices safely, and ensure robust cost calculation:

--- a/litellm/main.py
+++ b/litellm/main.py
@@ -8813,32 +8813,44 @@ def stream_chunk_builder(
         provider_specific_chunks = [
             chunk
             for chunk in chunks
-            if (
-                isinstance(chunk, dict)
-                and chunk.get("choices")
-                and isinstance(chunk["choices"], list)
-                and len(chunk["choices"]) > 0
-                and isinstance(chunk["choices"][0], dict)
-                and chunk["choices"][0].get("delta")
-                and isinstance(chunk["choices"][0]["delta"], dict)
-                and "provider_specific_fields" in chunk["choices"][0]["delta"]
-            )
+            if chunk is not None
         ]
 
         if len(provider_specific_chunks) > 0:
-            provider_field_dicts: Final = tuple(
-                fields
-                for chunk in provider_specific_chunks
-                for fields in (chunk["choices"][0]["delta"]["provider_specific_fields"],)
-                if isinstance(fields, dict)
-            )
+            collected_fields: list[dict[str, object]] = []
+            for chunk in provider_specific_chunks:
+                # Extract choices whether chunk is a dict or a ModelResponseStream
+                choices = chunk.get("choices", []) if isinstance(chunk, dict) else getattr(chunk, "choices", [])
+                if not choices:
+                    continue
+                for choice in choices:
+                    delta = choice.get("delta") if isinstance(choice, dict) else getattr(choice, "delta", None)
+                    if delta is None:
+                        continue
+                    fields = (
+                        delta.get("provider_specific_fields")
+                        if isinstance(delta, dict)
+                        else getattr(delta, "provider_specific_fields", None)
+                    )
+                    if isinstance(fields, dict):
+                        collected_fields.append(fields)
+
+            provider_field_dicts: Final = tuple(collected_fields)
             streamed_citations: Final = tuple(
                 fields["citation"] for fields in provider_field_dicts if fields.get("citation") is not None
             )
             citation_fields: Final = (
                 {"citations": [list(streamed_citations)]} if streamed_citations else {}  # mutable-ok: JSON dict field
             )
             combined_provider_fields: Final = {  # mutable-ok: Message.provider_specific_fields is a plain dict field
                 key: value
                 for fields in (citation_fields, *provider_field_dicts)
                 for key, value in fields.items()
                 if key != "citation"
             }

5. Verification & Test Suite Recommendations

Missing Test Cases to Add:

  1. test_stream_chunk_builder_citations_with_model_response_stream_objects: Test with ModelResponseStream instances to verify Pydantic object traversal.
  2. test_stream_chunk_builder_citations_multi_choice: Test with $n=2$ streamed completions.
  3. test_stream_chunk_builder_empty_delta_keepalive: Test with SSE keep-alive chunks where choices: [].

6. Final Review Verdict

Verdict: REQUEST_CHANGES
Rationale: While the citation aggregation logic is sound for dictionary-based chunk mocks, it completely fails in production when supplied with native ModelResponseStream objects due to strict isinstance(chunk, dict) checks. Implementing the polymorphic extractor diff above resolves all failure modes.

@codspeed

codspeed Bot commented Aug 28, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_lit6376_dspy_streaming (9b1b8e7) with litellm_internal_staging (b724ebc)

Open in CodSpeed

@mateo-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

mateo-berri commented Aug 28, 2026 •

Copy link
Copy Markdown
Contributor Author

@dlowzzxx This seems like an auto-generated review. All of that is pre-existing and not modified by this PR

@mateo-berri mateo-berri added run-ci and removed run-ci labels Aug 28, 2026
@mateo-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

Comment thread litellm/main.py
if isinstance(usage_cost, (int, float)):
return float(usage_cost)
if logging_obj is not None:
return None

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stamped stream cost skips corrected model

Medium Severity

_stream_builder_response_cost still copies usage.cost into _hidden_params["response_cost"] when a logging_obj is present. Mid-stream failure recovery can set usage.cost via _response_cost_calculator before it overwrites partial_response.model, then _response_cost_calculator trusts that stamped value and never re-prices with the corrected model, custom pricing, or later logging-pipeline adjustments.

Additional Locations (2)
Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit 9b1b8e7. Configure here.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Provider-reported usage.cost wins by design: response_cost_calculator returns it before custom pricing on the success path, so failure recovery pricing matches a completed stream

@mateo-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

1 issue from previous review remains unresolved.

Fix All in Cursor

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 9b1b8e7. Configure here.

@cursor

cursor Bot commented Aug 28, 2026

Copy link
Copy Markdown
Contributor

Bugbot Autofix prepared a fix for the issue found in the latest run.

  • ✅ Fixed: Stamped stream cost skips corrected model
    • Reordered _stream_builder_response_cost so passing a logging_obj skips the _hidden_params["response_cost"] stamp entirely, letting failure recovery reprice with the corrected model.

Create PR

Or push these changes by commenting:

@cursor push 41b9a41667
Preview (41b9a41667)
diff --git a/litellm/main.py b/litellm/main.py
--- a/litellm/main.py
+++ b/litellm/main.py
@@ -8591,11 +8591,11 @@
 
 
 def _stream_builder_response_cost(response: ModelResponse, logging_obj: Optional["Logging"]) -> float | None:
+    if logging_obj is not None:
+        return None
     usage_cost: Final = getattr(getattr(response, "usage", None), "cost", None)
     if isinstance(usage_cost, (int, float)):
         return float(usage_cost)
-    if logging_obj is not None:
-        return None
     provider_hint: Final = response._hidden_params.get(  # pyright: ignore[reportPrivateUsage]  # no public accessor
         "custom_llm_provider"
     )

diff --git a/tests/test_litellm/litellm_core_utils/test_streaming_handler.py b/tests/test_litellm/litellm_core_utils/test_streaming_handler.py
--- a/tests/test_litellm/litellm_core_utils/test_streaming_handler.py
+++ b/tests/test_litellm/litellm_core_utils/test_streaming_handler.py
@@ -3484,6 +3484,52 @@
     assert logging_obj.model_call_details["response_cost"] == pytest.approx(expected)
 
 
+def test_stream_chunk_builder_defers_response_cost_when_logging_obj_present(monkeypatch):
+    """Regression: passing logging_obj into stream_chunk_builder must leave the
+    downstream logging pipeline in charge of pricing. Stamping response_cost
+    on _hidden_params from the raw chunk model would short-circuit the
+    logging-object recalculator and bake the wrong price into failure spend.
+    """
+    monkeypatch.setattr(litellm, "include_cost_in_streaming_usage", True)
+    logging_obj = Logging(
+        model="gpt-4o-mini",
+        messages=[{"role": "user", "content": "hi"}],
+        stream=True,
+        call_type="completion",
+        start_time=time.time(),
+        litellm_call_id="stamped-cost-defer",
+        function_id="1245",
+    )
+    logging_obj.model_call_details["custom_llm_provider"] = "openai"
+    logging_obj.optional_params = {}
+
+    chunks = [
+        ModelResponseStream(
+            id="chatcmpl-defer-1",
+            created=1742056047,
+            model="claude-opus-4-5",
+            object="chat.completion.chunk",
+            choices=[
+                StreamingChoices(
+                    finish_reason=None,
+                    index=0,
+                    delta=Delta(content="hi", role="assistant"),
+                )
+            ],
+            usage=Usage(prompt_tokens=40, completion_tokens=5, total_tokens=45),
+        )
+    ]
+
+    partial = litellm.stream_chunk_builder(
+        chunks=chunks,
+        messages=[{"role": "user", "content": "hi"}],
+        logging_obj=logging_obj,
+    )
+
+    assert partial is not None
+    assert "response_cost" not in partial._hidden_params
+
+
 def test_record_partial_usage_for_failure_carries_up_openai_style_cached_tokens():
     recovered = Usage(
         prompt_tokens=1000,

You can send follow-ups to the cloud agent here.

@mateo-berri
mateo-berri enabled auto-merge August 28, 2026 22:28
@mateo-berri
mateo-berri merged commit 4ebedf9 into litellm_internal_staging Aug 28, 2026
129 of 131 checks passed
@mateo-berri
mateo-berri deleted the litellm_lit6376_dspy_streaming branch August 28, 2026 22:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants