fix(exceptions): keep internal_server_error as the public type of an upstream 500 - #41930
Conversation
…upstream 500 PR #40243 started carrying the upstream error body on InternalServerError so the Responses response.failed event can report the provider's code and message, and openai's APIError.__init__ took the body's type along with it. The proxy then answered an OpenAI-compatible upstream 500 with type server_error while a 502 and a 503 kept internal_server_error, and the integration contract in test_observed_routing.py went red. Pin the type the way RateLimitError pins throttling_error, keeping the body.
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
|
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 5f6ffdc. Configure here.
TLDR
Problem this solves:
type: server_errorsince fix(responses): emit typed streaming failure events #40243internal_server_errorintegration-providerson main is red at that assertionHow it solves it:
InternalServerErrorpinstypetointernal_server_error, likeRateLimitErrorpinsthrottling_errorresponse.failedeventUser Flow
Before: a developer whose client branches on
error.typesees a provider 500 labelled differently from a provider 502 or 503{"model": "gpt-5.4-mini", "messages": [{"role": "user", "content": "hi"}]}while the provider answers 500{"error": {"message": "litellm.InternalServerError: InternalServerError: OpenAIException - Controlled provider failure ...", "type": "server_error", "param": null, "code": "500"}}"type": "internal_server_error"internal_server_errorcatches the 503 and misses the 500After: the provider 500 carries the same
internal_server_errorlabel as the 502 and 503{"model": "gpt-5.4-mini", "messages": [{"role": "user", "content": "hi"}]}while the provider answers 500{"error": {"message": "litellm.InternalServerError: InternalServerError: OpenAIException - Controlled provider failure ...", "type": "internal_server_error", "param": null, "code": "500"}}"type": "internal_server_error"internal_server_errorcatches bothRelevant issues
Regression from #40243. The contract it broke landed two days earlier in #41075
Affected release
Linear ticket
Why the mapping changes and the assertion stays
#40243 set out to give the Responses
response.failedevent the upstream's code and message, so it started carrying the upstream body onInternalServerError. Thetypechange came along for the ride: openai'sAPIError.__init__copiestypeout of the body, and the proxy's error payload lets a carried type win over the status-derived one. It appears only in #40243's Low caveats, nobody discussed it as a contract change, and it lands on exactly one status, since the 502 and 503 mappers never receive a body andRateLimitErroralready pins its own label. The integration contract in #41075 pinnedinternal_server_errorfor a 500. So this PR treats the echo as a byproduct, restores the label, and keeps the bodyOpen decision, not taken here: whether the proxy should relay the upstream's own label for every 5xx (OpenAI itself labels a 500
server_error) uniformly across 500, 502, and 503. That would be a deliberate contract change with its own tests and docs rather than a side effect of one mapperThe proxy path itself is guarded by the integration contract this PR turns green (
tests/integration/routing/test_observed_routing.py, theintegration-providersjob); the two unit tests pin the exception and the payload helper underneath itPre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Two legs, one proxy instance each with
--num_workers 2(thelsofline at the end of each leg shows the master plus two worker pids listening), both booted the way.circleci/scripts/run_integration.shboots the CI integration rig: Postgres on 127.0.0.1:5432 (databasemat602_qa_<leg>),redis-server --port <redis>, the owned upstreampython -m integration._support.upstream --port <upstream>, and the proxypython -m integration._support.proxy --config tests/integration/proxy_config.yaml --host 127.0.0.1 --port <proxy> --num_workers 2 --telemetry False --use_prisma_db_push --enforce_prisma_migration_checkwithLITELLM_MODE=PRODUCTION STORE_MODEL_IN_DB=True LITELLM_LOCAL_MODEL_COST_MAP=True LITELLM_MASTER_KEY=sk-integration-master, readiness through.circleci/scripts/wait_integration_services.py, then the CI seedPOST /config/update {"router_settings": {"num_retries": 0}}. Before ran faed57f (the merge base) on proxy 26506, upstream 48335, redis 47462; After ran 5f6ffdc (this PR's tip) on proxy 22096, upstream 27501, redis 51912Two deployments are registered per leg through
POST /model/newwithapi_basepointing at the upstream:mat602-<leg>isopenai/mat602-errors-<leg>, the provider the ticket names, andmat602-bridge-<leg>iscustom_openai/mat602-bridge-<leg>, needed because the owned upstream serves only the OpenAI chat completions route, so/v1/messagesand/v1/responseson it bridge to the upstream's/v1/chat/completionsthrough the same openai SDK path that carries the upstream body. Each request is scripted throughPOST /__scripts/<model> {"statuses": [<code>]}right before it and cleared after. The upstream's own answer to a 500, with no proxy in the path, is{"error":{"message":"Controlled provider failure","type":"server_error","code":"500"}}(control, identical on both legs), and its/__observationslisted every bridged request on/v1/chat/completions, one attempt eachBefore (faed57f)
The integration node the ticket names
A chat completion while the upstream answers 500
The same request while the upstream answers 502, 503, 429, and 400
The same chat completion with stream: true while the upstream answers 500
A chat completion on the custom_openai/ deployment while the upstream answers 500
/v1/messages on the custom_openai/ deployment while the upstream answers 500, plain and stream: true
The
"stream": truevariant answered the byte-identical body and HTTP 500/v1/responses on the custom_openai/ deployment while the upstream answers 500, plain and stream: true
The
"stream": truevariant answered the byte-identical body and HTTP 500After (5f6ffdc)
The integration node the ticket names
A chat completion while the upstream answers 500
The same request while the upstream answers 502, 503, 429, and 400
The same chat completion with stream: true while the upstream answers 500
A chat completion on the custom_openai/ deployment while the upstream answers 500
/v1/messages on the custom_openai/ deployment while the upstream answers 500, plain and stream: true
The
"stream": truevariant answered the byte-identical body and HTTP 500. The Anthropic-shaped error is built from the HTTP status alone, so it readsapi_erroron both legs/v1/responses on the custom_openai/ deployment while the upstream answers 500, plain and stream: true
The
"stream": truevariant answered the byte-identical body and HTTP 500The text completion route and the fallback header, both legs
A second pair of rigs, same shape (
--num_workers 2, master plus two workers on the proxy port), Before at faed57f on proxy 24603 and upstream 29417, After at 5f6ffdc on proxy 25119 and upstream 28874. The text completion route rewrites an unknownopenai/model to the provider's native/v1/completions, which the owned upstream cannot script, so the failing deployment here isopenai/gpt-4o(a model name LiteLLM knows as a chat model, which the route bridges onto the upstream's scripted/v1/chat/completions) under the aliasmat602-primary-<leg>, next to a healthyopenai/mat602-healthy-<leg>POST /v1/completions while the upstream answers 500, plain and stream: true
The
"stream": truevariant answered the byte-identical body and HTTP 500 on each leg:server_errorbefore,internal_server_errorafterPOST /v1/chat/completions with per-request fallbacks and include_fallback_errors while the primary's upstream answers 500
Not observed. With
POST /config/update {"general_settings": {"expose_fallback_errors_to_caller": true}}accepted ({"message":"Config updated successfully"}) and the request carrying"fallbacks": ["mat602-healthy-<leg>"], "include_fallback_errors": true, the fallback itself ran on both legs (HTTP 200 answered by the healthy deployment,x-litellm-attempted-fallbacks: 1, the upstream's/__observationsshowinggpt-4othenmat602-healthy-<leg>), but neither leg ever returned anx-litellm-fallback-errorsheader in 120 attempts over two minutes, before and after alike. The flag set at runtime through/config/updateis the same on both legs and outside this PR; a yaml-configured flag was not triedObserved next to the fix, on both legs alike:
internal_server_error; only 500 differed; fixed herethrottling_error; the integration test expects it; left alone/v1/messageserrors readapi_erroron both legs; status-derived, by designexpose_fallback_errors_to_callernever yields the header; pre-existing, untouchedGates:
make checkPASS at 5f6ffdc and at ddac683; the two touched test files plustest_exception_header_preservation.pyandproxy/test_common_request_processing.pypass locally at ddac683 (1648 passed), and 5f6ffdc differs from it only by a removed docstring in a test file. Both new assertions fail without the one-line fix (assert 'server_error' == 'internal_server_error')Type
🐛 Bug Fix
Caveats (if any)
Low
APIErrorstill copies a carried body'stype, unchanged here: no mapper builds one with a body today, and the litellm_proxy relay of a downstreamlitellm.APIErrorfails earlier on its own, both pre-existing.typeon a caughtlitellm.InternalServerErrorreadsinternal_server_errornow, where it wasNonebefore fix(responses): emit typed streaming failure events #40243 and the upstream's label after itx-litellm-fallback-errorsheader serializes that.type, so for a 500 it carriedNonebefore fix(responses): emit typed streaming failure events #40243 andserver_errorafter it; it carriesinternal_server_errornow. That is read off the header builder'sgetattr(error, "type", ...)and its unit tests, not observed live: withexpose_fallback_errors_to_callerset through/config/update, the header never appeared on either leg (fallback ran,x-litellm-attempted-fallbacks: 1), before and after alike.circleci/scripts/run_integration.shdoes and everything from the proxy outward is reallitellm_utils_testing::test_models_by_provider,proxy_e2e_anthropic_messages_tests::test_bedrock_invoke_messages_with_all_beta_headers[*]) and fail identically onmain(pipeline 89818, before this branch); neither touches this diff, and theintegrationworkflow the ticket names is green at the tipFinal Attestation
The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR
5f6ffdc passes /live-pr-risk