Skip to content

fix(proxy): retry rate-limit fallbacks from a pristine request snapshot - #40596

Merged
yucheng-berri merged 8 commits into
mainfrom
litellm_lit_7470_rate_limit_fallback_pristine_data
Sep 16, 2026
Merged

yucheng-berri merged 8 commits into
mainfrom
litellm_lit_7470_rate_limit_fallback_pristine_data

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Rate limited request with fallbacks configured returned HTTP 500 instead of 429
  • Error text was cannot pickle '_thread.RLock' object
  • Only happens with OpenTelemetry callbacks on, since the span lives in request metadata
  • Hit on every worker and every pod, not a single-instance quirk

How it solves it:

  • Snapshot the client request before the first pre-call pass, only when fallbacks are configured
  • Each fallback attempt starts from a fresh copy of that snapshot with the model swapped
  • Proxy-internal state added by the failed attempt (span, auth object, logging object) is never reprocessed
  • Fallback lookup still runs after alias normalization, so aliased models keep their fallbacks
  • On exhaustion or failure, the request is restored to the first attempt so the 429 names the requested model
  • disable_fallbacks from key metadata is checked after the first pass, since that pass is what writes it

User Flow

Before: a developer whose key is rate limited on a model that has a fallback gets an opaque 500

  1. The proxy admin runs with otel in litellm_settings.callbacks and fallbacks: [{"gpt-main": ["gpt-fallback"]}] in router_settings
  2. The admin creates a key with model_rpm_limit: {"gpt-main": 1}
  3. The developer sends POST https://litellm-domain/v1/chat/completions with "model": "gpt-main" and gets 200 from gpt-main
  4. They send the same request again within the minute
  5. Response is HTTP 500 with "cannot pickle '_thread.RLock' object", no retry-after header, no fallback

After: the same request lands on the fallback, or returns a real 429 when the fallback is limited too

  1. The proxy admin runs with otel in litellm_settings.callbacks and fallbacks: [{"gpt-main": ["gpt-fallback"]}] in router_settings
  2. The admin creates a key with model_rpm_limit: {"gpt-main": 1}
  3. The developer sends POST https://litellm-domain/v1/chat/completions with "model": "gpt-main" and gets 200 from gpt-main
  4. They send the same request again within the minute
  5. Response is HTTP 200 with x-litellm-model-group: gpt-fallback. If the fallback is limited too, the response is HTTP 429 with retry-after: 60 and a message naming the gpt-main limit

Relevant issues

Reported by a customer (Pylon #8439)

Affected release

regression in v1.93.0.dev1 (last working v1.92.2, still present through v1.101.0 and the v1.102.0-rc.1 line)

Linear ticket

Resolves LIT-7470

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Live audit of the dynamic_rate_limiter_v3 path, one leg per side, each from its own worktree with python -m litellm.proxy.proxy_cli --config config.yaml --port <port> --num_workers 2 --use_v2_migration_resolver, its own Postgres cluster and Redis container, real OpenAI and Anthropic calls, no mocks. Before ran the merge base a8979fe, After ran the PR tip 9974cf4, 127 graded cells in total across the three endpoints, streaming and non streaming, curl plus the OpenAI and Anthropic Python SDKs. Every 200 was matched to exactly one LiteLLM_SpendLogs row by litellm_call_id. Dashboard screenshots and the annotated recording of both legs are in this comment; the full matrix with every command, response id and spend row is in the audit report attached to the Devin session

model_list:
  - model_name: gpt-main
    litellm_params: {model: openai/gpt-5.4-nano, api_key: os.environ/OPENAI_API_KEY, rpm: 1}
  - model_name: gpt-fallback
    litellm_params: {model: openai/gpt-5.4-mini, api_key: os.environ/OPENAI_API_KEY}
  - model_name: gpt-ex
    litellm_params: {model: openai/gpt-5.4-nano, api_key: os.environ/OPENAI_API_KEY, rpm: 1}
  - model_name: gpt-ex-fb
    litellm_params: {model: openai/gpt-5.4-mini, api_key: os.environ/OPENAI_API_KEY, rpm: 1}
router_settings:
  fallbacks:
    - gpt-main: ["gpt-fallback"]
    - gpt-ex: ["gpt-ex-fb"]
litellm_settings:
  callbacks: ["otel", "dynamic_rate_limiter_v3"]

The limiter counts one request per pre-call, so with rpm: 1 the first request in the minute is served by the primary and every later one in that window is rejected locally. Each case below starts with that warm request. Helpers ($U is the leg URL, $J is Content-Type: application/json, show prints status, x-litellm-model-group, retry-after, id and error message)

key()  { curl -s "$U/key/generate" -H "Authorization: Bearer $MASTER_KEY" -H "$J" -d "${1:-{\}}" | python3 -c 'import sys,json;print(json.load(sys.stdin)["key"])'; }
chat() { curl -s -i "$U/v1/chat/completions" -H "Authorization: Bearer $1" -H "$J" -d "{\"model\":\"${3:-gpt-main}\",${2:-}\"messages\":[{\"role\":\"user\",\"content\":\"Say hi in 3 words\"}]}" | show; }
msgs() { curl -s -i "$U/v1/messages" -H "Authorization: Bearer $1" -H "$J" -d "{\"model\":\"${3:-gpt-main}\",${2:-}\"max_tokens\":64,\"messages\":[{\"role\":\"user\",\"content\":\"Say hi in 3 words\"}]}" | show; }
resp() { curl -s -i "$U/v1/responses" -H "Authorization: Bearer $1" -H "$J" -d "{\"model\":\"${3:-gpt-main}\",${2:-}\"input\":\"Say hi in 3 words\"}" | show; }

Before (a8979fe)

/v1/chat/completions, gpt-main saturated, fallback configured

  1. K=$(key), then chat $K returns HTTP 200 group=gpt-main id=chatcmpl-EOpUdRQAaDrMSaDjEGM2pLxZHNXWg
  2. chat $K returns HTTP 500 err=cannot pickle '_thread.RLock' object (call id 29dbc126-cb51-4d4f-bd1e-6de803acc0d7)
  3. chat $K '"stream":true,' returns HTTP 500 err=cannot pickle '_thread.RLock' object (call id 5234dedf-2215-46b9-96ce-d71b6cd2646f)
  4. OpenAI SDK, sync and async, streaming and not, against the same key: all four raise InternalServerError: 500 cannot pickle '_thread.RLock' object

/v1/messages, same key

  1. msgs $K returns HTTP 200 group=gpt-fallback (call id b57bf95e-f2ef-45ff-a8b2-428d0c7f9b8a; this route did not crash on base)
  2. msgs $K '"stream":true,' returns HTTP 200 group=gpt-fallback id=msg_d2db43f8…, stream consumed to message_stop
  3. Anthropic SDK, sync non streaming plus sync and async streaming, against the same key: three HTTP 200 group=gpt-fallback responses

/v1/responses, same key

  1. resp $K returns HTTP 200 group=gpt-fallback (call id 760359f1-8232-4423-bee8-f60ea8d04104; this route did not crash on base)
  2. resp $K '"stream":true,' returns HTTP 200 group=gpt-fallback id=resp_LhMgYxLa…, stream consumed to response.completed
  3. OpenAI SDK responses, sync non streaming plus sync and async streaming: three HTTP 200 group=gpt-fallback responses, streams end with response.completed

/v1/chat/completions, aliased model (key model_group_alias alias-main -> gpt-main)

  1. K3=$(key '{"router_settings":{"model_group_alias":{"alias-main":"gpt-main"}}}'), then chat $K3 '' gpt-main returns HTTP 200 group=gpt-main
  2. chat $K3 '' alias-main returns HTTP 500 err=cannot pickle '_thread.RLock' object (call id 3e436131-e7ff-4eec-af2c-32d42de10ee7), streaming variant the same

/v1/chat/completions, primary and only fallback both saturated (gpt-ex -> gpt-ex-fb)

  1. chat $K '' gpt-ex returns HTTP 200 group=gpt-ex, chat $K '' gpt-ex-fb returns HTTP 200 group=gpt-ex-fb
  2. chat $K '' gpt-ex returns HTTP 500 err=cannot pickle '_thread.RLock' object (call id 592c8796-78f3-4c8c-af81-8f4df5ec1bde)

/v1/chat/completions, key metadata disable_fallbacks: true

  1. KD=$(key '{"metadata":{"disable_fallbacks":true}}'), gpt-main already saturated
  2. chat $KD returns HTTP 429 retry-after=60 err=Model capacity reached for gpt-main ... Model RPM: 1, Remaining: 0 (no fallback attempted, same as After)

36 request mixed burst (12 per endpoint, half streaming), gpt-main saturated

  1. python burst.py base c0_control returns 12 x HTTP 500 cannot pickle '_thread.RLock' object (every chat completions request) and 24 x HTTP 200 group=gpt-fallback (messages and responses)
  2. verify_burst.sh base c0_control reports ok_ids=24 missing=0 duplicated=0 in LiteLLM_SpendLogs

After (9974cf4)

/v1/chat/completions, gpt-main saturated, fallback configured

  1. K=$(key), then chat $K returns HTTP 200 group=gpt-main id=chatcmpl-EOpT0U8myTI5YIwtr2kZ1UfdA9XNz
  2. chat $K returns HTTP 200 group=gpt-fallback id=chatcmpl-EOpT1f0Rll6MhBq725vsPzQcN2BmL, one spend row for call id 7098c00b-62ea-4fa9-9a05-3c003e74bff8 with model_group=gpt-fallback
  3. chat $K '"stream":true,' returns HTTP 200 group=gpt-fallback id=chatcmpl-EOpT1a33wcbSeE7r1nKPFGl3cVIQw, stream ends with data: [DONE], one spend row
  4. OpenAI SDK, sync and async, streaming and not, against the same key: four HTTP 200 group=gpt-fallback responses (chatcmpl-EOpT8ADm…, chatcmpl-EOpT9AxS…, chatcmpl-EOpTAeCq…, chatcmpl-EOpTAfdu…), each with one spend row

/v1/messages, same key

  1. msgs $K returns HTTP 200 group=gpt-fallback (call id 122a85cf-c41a-4244-85f3-7d2502155ddf), one spend row
  2. msgs $K '"stream":true,' returns HTTP 200 group=gpt-fallback id=msg_292b0698…, stream consumed to message_stop, one spend row
  3. Anthropic SDK, sync non streaming plus sync and async streaming (msg_879209d4…, msg_582010fc…): three HTTP 200 group=gpt-fallback responses, each with one spend row

/v1/responses, same key

  1. resp $K returns HTTP 200 group=gpt-fallback (call id f34fbf66-09e6-4783-ae8a-e180e024ec04), one spend row
  2. resp $K '"stream":true,' returns HTTP 200 group=gpt-fallback id=resp_6K8-T-tn…, stream consumed to response.completed, one spend row
  3. OpenAI SDK responses, sync non streaming plus sync and async streaming: three HTTP 200 group=gpt-fallback responses, streams end with response.completed, each with one spend row

/v1/chat/completions, aliased model (key model_group_alias alias-main -> gpt-main)

  1. K3=$(key '{"router_settings":{"model_group_alias":{"alias-main":"gpt-main"}}}'), then chat $K3 '' gpt-main returns HTTP 200 group=gpt-main
  2. chat $K3 '' alias-main returns HTTP 200 group=gpt-fallback id=chatcmpl-EOpXEdOY7xZJ95GQuzPJDCkj0GCgB, spend row shows the request served by gpt-fallback

/v1/chat/completions, primary and only fallback both saturated (gpt-ex -> gpt-ex-fb)

  1. chat $K '' gpt-ex returns HTTP 200 group=gpt-ex, chat $K '' gpt-ex-fb returns HTTP 200 group=gpt-ex-fb
  2. chat $K '' gpt-ex returns HTTP 429 retry-after=60 err=Model capacity reached for gpt-ex ... Model RPM: 1, Remaining: 0, and the streaming variant returns the same 429

/v1/chat/completions, key metadata disable_fallbacks: true

  1. KD=$(key '{"metadata":{"disable_fallbacks":true}}'), gpt-main already saturated
  2. chat $KD returns HTTP 429 retry-after=60 err=Model capacity reached for gpt-main ... Model RPM: 1, Remaining: 0, no fallback attempted

36 request mixed burst (12 per endpoint, half streaming), gpt-main saturated

  1. python burst.py head c1_redis with docker stop of the leg's Redis 1.5 s in and docker start at 12 s: 36 x HTTP 200 group=gpt-fallback, verify_burst.sh reports ok_ids=36 missing=0 duplicated=0
  2. python burst.py head c2_pgpause with kill -STOP on the leg's Postgres cluster 1.5 s in and kill -CONT at 15 s: 36 x HTTP 200 group=gpt-fallback, /health/readiness showed "db":"disconnected" during the pause, after resume ok_ids=36 missing=0 duplicated=0
  3. python burst.py head c4_restart with kill -TERM of the proxy parent 1.5 s in and a relaunch from the same tree: 36 x HTTP 200 group=gpt-fallback, ok_ids=36 missing=0 duplicated=0
  4. python burst.py head c3_workerkill with kill -9 of one uvicorn worker 1.5 s in: the sibling worker kept serving and uvicorn respawned the killed one, 20 x HTTP 200 group=gpt-fallback, 16 transport disconnects for requests in flight on the killed worker, and 5 of the 20 successful ids never got a spend row because the killed worker's in-memory spend queue died with it. This is what SIGKILL does to a uvicorn worker on both sides and is unrelated to the fix; recorded as a limitation, not a pass

Type

🐛 Bug Fix

Caveats (if any)

Low

  • With any fallbacks configured, every request pays one snapshot copy before the first pre-call pass, even for models without a fallback entry. Strings are shared by deepcopy, so the cost scales with container count, not payload size
  • A fallback attempt that raises something other than a rate limit error still propagates as before (existing test pins this)
  • With OTel off, base and head return identical statuses and bodies for fallback success, exhausted fallback, key-wide limit, disable_fallbacks: true in the body or in key metadata, and a key-level fallback pointing at a missing model group (400 on both)
  • A key-level router_settings.fallbacks: [] means no fallbacks (429) on both base and head; only null or an absent key falls through to the global list
  • A fallback retry still bypasses the key's models allowlist unless enforce_fallback_model_access is set, on both base and head; predates this PR, follow-up ticket material
  • Veria posted no review or check run on this PR at the tip, so that reviewer is recorded as unavailable rather than passing. Greptile (5/5) and Bugbot (no new issues) both reviewed 9974cf4
  • Legacy limiter (LEGACY_MULTI_INSTANCE_RATE_LIMITING=true) was not driven live. It raises the same error through the same seam and the retry code does not inspect the limiter, so the behavior should match, but this is inferred, not observed

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

ran /live-pr-risk and found no regressions/backward incompatible risks

Link to Devin session: https://app.devin.ai/sessions/cdbe8de849c546a8a5c1c0f9e635ced3
Open in Devin Desktop: https://app.devin.ai/desktop/session/cdbe8de849c546a8a5c1c0f9e635ced3?variant=devin
Requested by: @yucheng-berri

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration
devin-ai-integration Bot requested a review from a team September 10, 2026 17:17
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

CLAassistant commented Sep 10, 2026 •

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you all sign our Contributor License Agreement before we can accept your contribution.
1 out of 2 committers have signed the CLA.

✅ yucheng-berri
❌ devin-ai-integration[bot]
You have signed the CLA already but the status is still pending? Let us recheck it.

@codspeed

codspeed Bot commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_lit_7470_rate_limit_fallback_pristine_data (9974cf4) with main (5960881)1

Open in CodSpeed

Footnotes

  1. No successful run was found on main (a8979fe) during the generation of this report, so 5960881 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

@codecov

codecov Bot commented Sep 10, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 95.00000% with 1 line in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
litellm/proxy/common_request_processing.py 95.00% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@greptile-apps

greptile-apps Bot commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR fixes rate-limit fallback retries by preserving a pristine request snapshot before pre-call processing and rebuilding each fallback attempt from that snapshot. It retains alias-resolved fallback lookup, honors key-level fallback disabling, and restores the primary attempt’s state when retries fail or are exhausted. Regression tests cover OpenTelemetry metadata, successful and exhausted fallbacks, model aliases, and key metadata behavior

Confidence Score: 5/5

The PR appears safe to merge, with no outstanding correctness, security, or repository-rule violations

The alias fallback issue is fixed by resolving fallback models after the first pre-call pass, and the previously disputed test-mocking finding was withdrawn after confirming the production path remains intact. The test-helper typing thread was manually resolved without explanation. No changes were made after the previous review, and no new issues were identified

Important Files Changed
Filename Overview
litellm/proxy/common_request_processing.py Snapshots pristine request data before pre-call enrichment and reconstructs each configured fallback attempt without carrying non-copyable internal state
tests/test_litellm/proxy/test_common_request_processing.py Adds typed regression coverage using the real v3 limiter path for OpenTelemetry metadata, aliases, fallback exhaustion, and disabled fallbacks
tests/proxy_unit_tests/test_response_polling_pre_call_checks.py Uses a concrete UserAPIKeyAuth instance so fallback configuration inspection follows the production data contract

Reviews (5): Last reviewed commit: "fix(proxy): honor key-level disable_fall..." | Re-trigger Greptile

Comment thread litellm/proxy/common_request_processing.py Outdated
Comment thread tests/test_litellm/proxy/test_common_request_processing.py Outdated
…n fallback resolution

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@yuneng-berri
yuneng-berri deleted the branch main September 13, 2026 04:51
@yuneng-berri yuneng-berri reopened this Sep 13, 2026
yucheng-berri and others added 3 commits September 16, 2026 06:57
…d retry from a client-request snapshot

The fallback retry in _pre_call_with_fallbacks re-entered common_processing_pre_call_logic with data already enriched by the first pass, so add_litellm_data_to_request deep-copied a metadata dict holding the live OTel span and the request failed with a 500 (cannot pickle '_thread.RLock') instead of the intended 429 or fallback. Capture the configured fallbacks and a snapshot of the client request before the first pass, look up the fallback chain by the normalized model group after the limiter raises, and run each fallback attempt on a fresh copy of that snapshot. Replaces the mock-heavy tests with a rig that runs the real v3 limiter and a live OTel span through the proxy_logging_obj seam

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ead of building a dict literal

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration
devin-ai-integration Bot changed the base branch from litellm_internal_staging to main September 16, 2026 07:45
@devin-ai-integration devin-ai-integration Bot added the backport-stable P0 regression fix only (Urgent ticket): cherry-pick onto the baking rc line before the stable tag label Sep 16, 2026
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread tests/test_litellm/proxy/test_common_request_processing.py Outdated
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/proxy/common_request_processing.py Outdated
Key metadata disable_fallbacks only lands on data during add_key_level_controls,
so the local rate-limit fallback retry now rechecks it post pre-call. Also use a
real UserAPIKeyAuth in the skip pre-call test since the path reads router_settings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 9974cf4. Configure here.

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

Live audit on dynamic_rate_limiter_v3: base a8979fe returns the RLock 500, head 9974cf4 serves the fallback. Full report in the session

Base a8979fe, Playground second send Head 9974cf4, Playground second send
Base Playground second send fails with the RLock 500 Head Playground second send answered by the fallback
Base Logs, failed request detail Head Logs, fallback request metadata
Base Logs detail shows cannot pickle RLock Head Logs metadata shows gpt-fallback served the aliased request
Head Settings, router fallbacks Head Settings, callbacks
Router settings show gpt-main falling back to gpt-fallback Callbacks show OpenTelemetry and dynamic_rate_limiter_v3
Annotated recording of the dashboard flow on both legs

Recording

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@yucheng-berri
yucheng-berri merged commit c5325b1 into main Sep 16, 2026
96 checks passed
@yucheng-berri
yucheng-berri deleted the litellm_lit_7470_rate_limit_fallback_pristine_data branch September 16, 2026 23:47
mateo-berri added a commit that referenced this pull request Sep 18, 2026
…rc_1_102_0

fix: backport eight backport-stable fixes to rc/1.102.0 (#40596, #41046, #41086, #41171, #41178, #41283, #41495, #41689)
jibanez-staticduo pushed a commit to jibanez-staticduo/litellm that referenced this pull request Sep 18, 2026
… on rate-limit fallback

The local rate-limit fallback path introduced in BerriAI#40596 restores the request from a snapshot taken before the pre-call pass, so the guardrails resolved for the requested model were dropped when another deployment was selected, and the raw request-body disable_fallbacks field gated the fallback before the key-level disable_fallbacks override had been applied

Carry the requested model's merged guardrail list onto every fallback pass and read the effective disable_fallbacks value after the pre-call pass

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
tech-carrement pushed a commit to tech-carrement/litellm that referenced this pull request Sep 20, 2026
Backport of BerriAI#40596 to rc/1.102.0.
Cherry-picked from merge commit c5325b1 (main), originally by app/devin-ai-integration.
@mateo-berri mateo-berri removed the backport-stable P0 regression fix only (Urgent ticket): cherry-pick onto the baking rc line before the stable tag label Sep 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants