Skip to content

fix(exceptions): keep internal_server_error as the public type of an upstream 500 - #41930

Merged
mateo-berri merged 2 commits into
mainfrom
litellm_mat602_upstream_500_error_type
Sep 19, 2026
Merged

mateo-berri merged 2 commits into
mainfrom
litellm_mat602_upstream_500_error_type

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

How it solves it:

  • InternalServerError pins type to internal_server_error, like RateLimitError pins throttling_error
  • The upstream body still rides along for the Responses response.failed event

User Flow

Before: a developer whose client branches on error.type sees a provider 500 labelled differently from a provider 502 or 503

  1. They send POST https://litellm-domain/v1/chat/completions with {"model": "gpt-5.4-mini", "messages": [{"role": "user", "content": "hi"}]} while the provider answers 500
  2. They get HTTP 500 with {"error": {"message": "litellm.InternalServerError: InternalServerError: OpenAIException - Controlled provider failure ...", "type": "server_error", "param": null, "code": "500"}}
  3. They send the same request while the provider answers 503 and get HTTP 503 with "type": "internal_server_error"
  4. Their handler for internal_server_error catches the 503 and misses the 500

After: the provider 500 carries the same internal_server_error label as the 502 and 503

  1. They send POST https://litellm-domain/v1/chat/completions with {"model": "gpt-5.4-mini", "messages": [{"role": "user", "content": "hi"}]} while the provider answers 500
  2. They get HTTP 500 with {"error": {"message": "litellm.InternalServerError: InternalServerError: OpenAIException - Controlled provider failure ...", "type": "internal_server_error", "param": null, "code": "500"}}
  3. They send the same request while the provider answers 503 and get HTTP 503 with "type": "internal_server_error"
  4. Their handler for internal_server_error catches both

Relevant issues

Regression from #40243. The contract it broke landed two days earlier in #41075

Affected release

Linear ticket

Why the mapping changes and the assertion stays

#40243 set out to give the Responses response.failed event the upstream's code and message, so it started carrying the upstream body on InternalServerError. The type change came along for the ride: openai's APIError.__init__ copies type out of the body, and the proxy's error payload lets a carried type win over the status-derived one. It appears only in #40243's Low caveats, nobody discussed it as a contract change, and it lands on exactly one status, since the 502 and 503 mappers never receive a body and RateLimitError already pins its own label. The integration contract in #41075 pinned internal_server_error for a 500. So this PR treats the echo as a byproduct, restores the label, and keeps the body

Open decision, not taken here: whether the proxy should relay the upstream's own label for every 5xx (OpenAI itself labels a 500 server_error) uniformly across 500, 502, and 503. That would be a deliberate contract change with its own tests and docs rather than a side effect of one mapper

The proxy path itself is guarded by the integration contract this PR turns green (tests/integration/routing/test_observed_routing.py, the integration-providers job); the two unit tests pin the exception and the payload helper underneath it

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Two legs, one proxy instance each with --num_workers 2 (the lsof line at the end of each leg shows the master plus two worker pids listening), both booted the way .circleci/scripts/run_integration.sh boots the CI integration rig: Postgres on 127.0.0.1:5432 (database mat602_qa_<leg>), redis-server --port <redis>, the owned upstream python -m integration._support.upstream --port <upstream>, and the proxy python -m integration._support.proxy --config tests/integration/proxy_config.yaml --host 127.0.0.1 --port <proxy> --num_workers 2 --telemetry False --use_prisma_db_push --enforce_prisma_migration_check with LITELLM_MODE=PRODUCTION STORE_MODEL_IN_DB=True LITELLM_LOCAL_MODEL_COST_MAP=True LITELLM_MASTER_KEY=sk-integration-master, readiness through .circleci/scripts/wait_integration_services.py, then the CI seed POST /config/update {"router_settings": {"num_retries": 0}}. Before ran faed57f (the merge base) on proxy 26506, upstream 48335, redis 47462; After ran 5f6ffdc (this PR's tip) on proxy 22096, upstream 27501, redis 51912

Two deployments are registered per leg through POST /model/new with api_base pointing at the upstream: mat602-<leg> is openai/mat602-errors-<leg>, the provider the ticket names, and mat602-bridge-<leg> is custom_openai/mat602-bridge-<leg>, needed because the owned upstream serves only the OpenAI chat completions route, so /v1/messages and /v1/responses on it bridge to the upstream's /v1/chat/completions through the same openai SDK path that carries the upstream body. Each request is scripted through POST /__scripts/<model> {"statuses": [<code>]} right before it and cleared after. The upstream's own answer to a 500, with no proxy in the path, is {"error":{"message":"Controlled provider failure","type":"server_error","code":"500"}} (control, identical on both legs), and its /__observations listed every bridged request on /v1/chat/completions, one attempt each

Before (faed57f)

The integration node the ticket names

$ .venv/bin/python -m pytest 'tests/integration/routing/test_observed_routing.py::test_retry_counts_and_public_errors_match_actual_provider_attempts' -vv
faed57f92c Merge pull request #41918 from BerriAI/litellm_websearch_followup_api_base
tests/integration/routing/test_observed_routing.py::test_retry_counts_and_public_errors_match_actual_provider_attempts FAILED [100%]
>                       assert error["type"] == {400: "invalid_request_error", 429: "throttling_error", 500: "internal_server_error"}[status]
E                       AssertionError: assert 'server_error' == 'internal_server_error'
tests/integration/routing/test_observed_routing.py:45: AssertionError
============================== 1 failed in 9.08s ===============================

A chat completion while the upstream answers 500

$ curl -X POST http://127.0.0.1:26506/v1/chat/completions -H 'Authorization: Bearer $LITELLM_MASTER_KEY' -d '{"model": "mat602-before", "messages": [{"role": "user", "content": "hi"}]}'
{"error":{"message":"litellm.InternalServerError: InternalServerError: OpenAIException - Controlled provider failure. Received Model Group=mat602-before\nAvailable Model Group Fallbacks=None","type":"server_error","param":null,"code":"500"}}
HTTP 500

The same request while the upstream answers 502, 503, 429, and 400

HTTP 502  "type":"internal_server_error"
HTTP 503  "type":"internal_server_error"
HTTP 429  "type":"throttling_error"
HTTP 400  "type":"invalid_request_error"

The same chat completion with stream: true while the upstream answers 500

$ curl -X POST http://127.0.0.1:26506/v1/chat/completions -H 'Authorization: Bearer $LITELLM_MASTER_KEY' -d '{"model": "mat602-before", "stream": true, "messages": [{"role": "user", "content": "hi"}]}'
{"error":{"message":"litellm.InternalServerError: InternalServerError: OpenAIException - Controlled provider failure. Received Model Group=mat602-before\nAvailable Model Group Fallbacks=None","type":"server_error","param":null,"code":"500"}}
HTTP 500

A chat completion on the custom_openai/ deployment while the upstream answers 500

$ curl -X POST http://127.0.0.1:26506/v1/chat/completions -H 'Authorization: Bearer $LITELLM_MASTER_KEY' -d '{"model": "mat602-bridge-before", "messages": [{"role": "user", "content": "hi"}]}'
{"error":{"message":"litellm.InternalServerError: InternalServerError: Custom_openaiException - Controlled provider failure. Received Model Group=mat602-bridge-before\nAvailable Model Group Fallbacks=None","type":"server_error","param":null,"code":"500"}}
HTTP 500

/v1/messages on the custom_openai/ deployment while the upstream answers 500, plain and stream: true

$ curl -X POST http://127.0.0.1:26506/v1/messages -H 'Authorization: Bearer $LITELLM_MASTER_KEY' -d '{"model": "mat602-bridge-before", "max_tokens": 16, "messages": [{"role": "user", "content": "hi"}]}'
{"type":"error","error":{"type":"api_error","message":"litellm.InternalServerError: InternalServerError: Custom_openaiException - Controlled provider failure. Received Model Group=mat602-bridge-before\nAvailable Model Group Fallbacks=None"}}
HTTP 500

The "stream": true variant answered the byte-identical body and HTTP 500

/v1/responses on the custom_openai/ deployment while the upstream answers 500, plain and stream: true

$ curl -X POST http://127.0.0.1:26506/v1/responses -H 'Authorization: Bearer $LITELLM_MASTER_KEY' -d '{"model": "mat602-bridge-before", "input": "hi"}'
{"error":{"message":"litellm.InternalServerError: InternalServerError: Custom_openaiException - Controlled provider failure. Received Model Group=mat602-bridge-before\nAvailable Model Group Fallbacks=None","type":"server_error","param":null,"code":"500"}}
HTTP 500

The "stream": true variant answered the byte-identical body and HTTP 500

After (5f6ffdc)

The integration node the ticket names

$ .venv/bin/python -m pytest 'tests/integration/routing/test_observed_routing.py::test_retry_counts_and_public_errors_match_actual_provider_attempts' -vv
5f6ffdc333 test: drop the docstring that restated the payload test's name
tests/integration/routing/test_observed_routing.py::test_retry_counts_and_public_errors_match_actual_provider_attempts PASSED [100%]
============================== 1 passed in 10.64s ==============================

A chat completion while the upstream answers 500

$ curl -X POST http://127.0.0.1:22096/v1/chat/completions -H 'Authorization: Bearer $LITELLM_MASTER_KEY' -d '{"model": "mat602-after", "messages": [{"role": "user", "content": "hi"}]}'
{"error":{"message":"litellm.InternalServerError: InternalServerError: OpenAIException - Controlled provider failure. Received Model Group=mat602-after\nAvailable Model Group Fallbacks=None","type":"internal_server_error","param":null,"code":"500"}}
HTTP 500

The same request while the upstream answers 502, 503, 429, and 400

HTTP 502  "type":"internal_server_error"
HTTP 503  "type":"internal_server_error"
HTTP 429  "type":"throttling_error"
HTTP 400  "type":"invalid_request_error"

The same chat completion with stream: true while the upstream answers 500

$ curl -X POST http://127.0.0.1:22096/v1/chat/completions -H 'Authorization: Bearer $LITELLM_MASTER_KEY' -d '{"model": "mat602-after", "stream": true, "messages": [{"role": "user", "content": "hi"}]}'
{"error":{"message":"litellm.InternalServerError: InternalServerError: OpenAIException - Controlled provider failure. Received Model Group=mat602-after\nAvailable Model Group Fallbacks=None","type":"internal_server_error","param":null,"code":"500"}}
HTTP 500

A chat completion on the custom_openai/ deployment while the upstream answers 500

$ curl -X POST http://127.0.0.1:22096/v1/chat/completions -H 'Authorization: Bearer $LITELLM_MASTER_KEY' -d '{"model": "mat602-bridge-after", "messages": [{"role": "user", "content": "hi"}]}'
{"error":{"message":"litellm.InternalServerError: InternalServerError: Custom_openaiException - Controlled provider failure. Received Model Group=mat602-bridge-after\nAvailable Model Group Fallbacks=None","type":"internal_server_error","param":null,"code":"500"}}
HTTP 500

/v1/messages on the custom_openai/ deployment while the upstream answers 500, plain and stream: true

$ curl -X POST http://127.0.0.1:22096/v1/messages -H 'Authorization: Bearer $LITELLM_MASTER_KEY' -d '{"model": "mat602-bridge-after", "max_tokens": 16, "messages": [{"role": "user", "content": "hi"}]}'
{"type":"error","error":{"type":"api_error","message":"litellm.InternalServerError: InternalServerError: Custom_openaiException - Controlled provider failure. Received Model Group=mat602-bridge-after\nAvailable Model Group Fallbacks=None"}}
HTTP 500

The "stream": true variant answered the byte-identical body and HTTP 500. The Anthropic-shaped error is built from the HTTP status alone, so it reads api_error on both legs

/v1/responses on the custom_openai/ deployment while the upstream answers 500, plain and stream: true

$ curl -X POST http://127.0.0.1:22096/v1/responses -H 'Authorization: Bearer $LITELLM_MASTER_KEY' -d '{"model": "mat602-bridge-after", "input": "hi"}'
{"error":{"message":"litellm.InternalServerError: InternalServerError: Custom_openaiException - Controlled provider failure. Received Model Group=mat602-bridge-after\nAvailable Model Group Fallbacks=None","type":"internal_server_error","param":null,"code":"500"}}
HTTP 500

The "stream": true variant answered the byte-identical body and HTTP 500

The text completion route and the fallback header, both legs

A second pair of rigs, same shape (--num_workers 2, master plus two workers on the proxy port), Before at faed57f on proxy 24603 and upstream 29417, After at 5f6ffdc on proxy 25119 and upstream 28874. The text completion route rewrites an unknown openai/ model to the provider's native /v1/completions, which the owned upstream cannot script, so the failing deployment here is openai/gpt-4o (a model name LiteLLM knows as a chat model, which the route bridges onto the upstream's scripted /v1/chat/completions) under the alias mat602-primary-<leg>, next to a healthy openai/mat602-healthy-<leg>

POST /v1/completions while the upstream answers 500, plain and stream: true

$ curl -X POST http://127.0.0.1:24603/v1/completions -H 'Authorization: Bearer $LITELLM_MASTER_KEY' -d '{"model": "mat602-primary-before2", "prompt": "hi"}'
{"error":{"message":"litellm.InternalServerError: InternalServerError: OpenAIException - Controlled provider failure. Received Model Group=mat602-primary-before2\nAvailable Model Group Fallbacks=None","type":"server_error","param":null,"code":"500"}}
HTTP 500

$ curl -X POST http://127.0.0.1:25119/v1/completions -H 'Authorization: Bearer $LITELLM_MASTER_KEY' -d '{"model": "mat602-primary-after2", "prompt": "hi"}'
{"error":{"message":"litellm.InternalServerError: InternalServerError: OpenAIException - Controlled provider failure. Received Model Group=mat602-primary-after2\nAvailable Model Group Fallbacks=None","type":"internal_server_error","param":null,"code":"500"}}
HTTP 500

The "stream": true variant answered the byte-identical body and HTTP 500 on each leg: server_error before, internal_server_error after

POST /v1/chat/completions with per-request fallbacks and include_fallback_errors while the primary's upstream answers 500

Not observed. With POST /config/update {"general_settings": {"expose_fallback_errors_to_caller": true}} accepted ({"message":"Config updated successfully"}) and the request carrying "fallbacks": ["mat602-healthy-<leg>"], "include_fallback_errors": true, the fallback itself ran on both legs (HTTP 200 answered by the healthy deployment, x-litellm-attempted-fallbacks: 1, the upstream's /__observations showing gpt-4o then mat602-healthy-<leg>), but neither leg ever returned an x-litellm-fallback-errors header in 120 attempts over two minutes, before and after alike. The flag set at runtime through /config/update is the same on both legs and outside this PR; a yaml-configured flag was not tried

Observed next to the fix, on both legs alike:

  • 502/503 already read internal_server_error; only 500 differed; fixed here
  • 429 reads throttling_error; the integration test expects it; left alone
  • Streaming errors arrive as plain JSON, not SSE; pre-existing, untouched
  • /v1/messages errors read api_error on both legs; status-derived, by design
  • Runtime-set expose_fallback_errors_to_caller never yields the header; pre-existing, untouched

Gates: make check PASS at 5f6ffdc and at ddac683; the two touched test files plus test_exception_header_preservation.py and proxy/test_common_request_processing.py pass locally at ddac683 (1648 passed), and 5f6ffdc differs from it only by a removed docstring in a test file. Both new assertions fail without the one-line fix (assert 'server_error' == 'internal_server_error')

Type

🐛 Bug Fix

Caveats (if any)

Low

  • APIError still copies a carried body's type, unchanged here: no mapper builds one with a body today, and the litellm_proxy relay of a downstream litellm.APIError fails earlier on its own, both pre-existing
  • The .type on a caught litellm.InternalServerError reads internal_server_error now, where it was None before fix(responses): emit typed streaming failure events #40243 and the upstream's label after it
  • The x-litellm-fallback-errors header serializes that .type, so for a 500 it carried None before fix(responses): emit typed streaming failure events #40243 and server_error after it; it carries internal_server_error now. That is read off the header builder's getattr(error, "type", ...) and its unit tests, not observed live: with expose_fallback_errors_to_caller set through /config/update, the header never appeared on either leg (fallback ran, x-litellm-attempted-fallbacks: 1), before and after alike
  • The proof's provider is the CI integration upstream, not a live vendor: no real provider answers a scripted 500 with a typed body on demand, and the ticket's requirement is that exact CI rig, so the legs boot it the way .circleci/scripts/run_integration.sh does and everything from the proxy outward is real
  • Two CircleCI jobs are red at the tip (litellm_utils_testing::test_models_by_provider, proxy_e2e_anthropic_messages_tests::test_bedrock_invoke_messages_with_all_beta_headers[*]) and fail identically on main (pipeline 89818, before this branch); neither touches this diff, and the integration workflow the ticket names is green at the tip

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

  • 5f6ffdc passes /live-pr-risk

…upstream 500

PR #40243 started carrying the upstream error body on InternalServerError so the Responses response.failed event can report the provider's code and message, and openai's APIError.__init__ took the body's type along with it. The proxy then answered an OpenAI-compatible upstream 500 with type server_error while a 502 and a 503 kept internal_server_error, and the integration contract in test_observed_routing.py went red. Pin the type the way RateLimitError pins throttling_error, keeping the body.
@devin-ai-integration

devin-ai-integration Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@greptile-apps

greptile-apps Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge; the public error-type correction is narrowly scoped and covered by focused regression tests.

Summary

Restores LiteLLM’s stable public error type for upstream HTTP 500 responses while preserving the upstream error body.

  • Pins InternalServerError.type to internal_server_error.
  • Adds coverage for exception mapping and proxy error payload generation.
  • Removes the redundant test docstring identified during the previous review.

Reviews (2) · Last reviewed commit: "test: drop the docstring that restated t..."

@codspeed

codspeed Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_mat602_upstream_500_error_type (5f6ffdc) with main (b7f0746)

Open in CodSpeed

Comment thread tests/test_litellm/proxy/common_utils/test_openai_error_payload.py Outdated
@codecov

codecov Bot commented Sep 19, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 5f6ffdc. Configure here.

@mateo-berri mateo-berri added run-ci and removed run-ci labels Sep 19, 2026

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@mateo-berri
mateo-berri merged commit 5611607 into main Sep 19, 2026
141 of 143 checks passed
@mateo-berri
mateo-berri deleted the litellm_mat602_upstream_500_error_type branch September 19, 2026 06:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant