feat(router): add capability classifier as Fuse foundation - #41270
Conversation
Codecov Report❌ Patch coverage is 📢 Thoughts on this report? Let us know! |
|
@greptileai Please review the updated capability classifier, including threshold routing, adaptive floors, configuration validation, logging controls, and dashboard preservation. |
Greptile SummaryAdds capability-based complexity routing as a foundation for Fuse:
Confidence Score: 5/5The PR appears safe to merge; no new actionable issue was introduced since the previous review, and the previous findings are resolved. The only change since the previous review reformats the encrypted-task opening-task conditional without changing its evaluation or output. Both previous findings are resolved, and no concrete mounted-rule violation remains.
|
| Filename | Overview |
|---|---|
| litellm/router_strategy/complexity_router/complexity_router.py | Integrates capability forecasting, fail-closed routing, adaptive floors, logging, and shared classifier calls; the post-review change is formatting-only. |
| litellm/router_strategy/complexity_router/capability_classifier.py | Defines the strict forecast contract, packaged capability prompt, verdict parsing, and threshold policy. |
| litellm/router_strategy/complexity_router/config.py | Adds strict capability, calibration, threshold, prompt-policy, and tier validation. |
| tests/test_litellm/router_strategy/test_complexity_router.py | Adds coverage for capability classification, calibration, failure behavior, task extraction, and routing integration. |
| ui/litellm-dashboard/src/components/add_model/build_complexity_router_config.ts | Preserves capability classifier configuration through dashboard configuration flows. |
Reviews (6): Last reviewed commit: "style(router): format encrypted task sel..." | Re-trigger Greptile
|
@greptileai Unknown capability policy fields now fail validation; the regression test reproduces the previous silent default and passes with the fix. |
|
bugbot run |
|
@greptileai Please review 896f35c, including request-scoped Codex task extraction and regressions covering generic callers, custom markers, and preserving original requests. |
|
@greptileai Please review cadb7ee after the Bugbot fixes for native encrypted task handling and the inherited classifier policy regressions |
|
bugbot run |
|
@greptileai Please refresh the review on 55fc0de; this commit only applies the formatter to the tested encrypted-task fix |
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 55fc0de. Configure here.
…103.0) (#290)
This PR contains the following updates:
| Package | Update | Change |
|---|---|---|
| [ghcr.io/berriai/litellm](https://images.chainguard.dev/directory/image/wolfi-base/overview) ([source](https://github.com/BerriAI/litellm)) | minor | `v1.102.1` → `v1.103.0` |
---
### Release Notes
<details>
<summary>BerriAI/litellm (ghcr.io/berriai/litellm)</summary>
### [`v1.103.0`](https://github.com/BerriAI/litellm/releases/tag/v1.103.0)
[Compare Source](https://github.com/BerriAI/litellm/compare/v1.102.1...v1.103.0)
#### Verify Docker Image Signature
All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0).
**Verify using the pinned commit hash (recommended):**
A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key:
```bash
cosign verify \
--key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \
ghcr.io/berriai/litellm:v1.103.0
```
**Verify using the release tag (convenience):**
Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules:
```bash
cosign verify \
--key https://raw.githubusercontent.com/BerriAI/litellm/v1.103.0/cosign.pub \
ghcr.io/berriai/litellm:v1.103.0
```
Expected output:
```
The following checks were performed on each of these signatures:
- The cosign claims were validated
- The signatures were verified against the specified public key
```
***
#### What's Changed
- fix(responses): translate the reasoning object into a chat-completion reasoning effort by [@​joshgarnett](https://github.com/joshgarnett) in [#​36363](https://github.com/BerriAI/litellm/pull/36363)
- fix(proxy): bound tool and guardrail index create\_many by the spend-log statement budgets by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40561](https://github.com/BerriAI/litellm/pull/40561)
- fix(mcp): require admission for delegated OAuth by [@​joshua-berri](https://github.com/joshua-berri) in [#​40923](https://github.com/BerriAI/litellm/pull/40923)
- fix(logging): log one bounded summary for a burst of timed-out LoggingWorker callbacks by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40912](https://github.com/BerriAI/litellm/pull/40912)
- fix(fireworks): resolve short model names to long cost map keys by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40929](https://github.com/BerriAI/litellm/pull/40929)
- ci: remove main branch source guard by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​40172](https://github.com/BerriAI/litellm/pull/40172)
- chore(ci): promote internal staging to main by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​40942](https://github.com/BerriAI/litellm/pull/40942)
- fix(spend\_logs): store litellm\_call\_id and match it in request\_id lookups by [@​mateo-berri](https://github.com/mateo-berri) in [#​39068](https://github.com/BerriAI/litellm/pull/39068)
- fix(auth): refresh lite login session token grants from the live user and team rows by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​40657](https://github.com/BerriAI/litellm/pull/40657)
- feat(bedrock): support file delete and list for S3-backed managed files by [@​mateo-berri](https://github.com/mateo-berri) in [#​39836](https://github.com/BerriAI/litellm/pull/39836)
- fix(proxy): gate the webhook test alert on proxy admins by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​40814](https://github.com/BerriAI/litellm/pull/40814)
- fix(ui): hide admin write-form tabs on the models page from view-only admins by [@​mateo-berri](https://github.com/mateo-berri) in [#​38867](https://github.com/BerriAI/litellm/pull/38867)
- fix(anthropic-adapter): surface mid-stream provider errors as Anthropic error events by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33352](https://github.com/BerriAI/litellm/pull/33352)
- docs(github): add an Affected release section to the PR template by [@​mateo-berri](https://github.com/mateo-berri) in [#​40618](https://github.com/BerriAI/litellm/pull/40618)
- docs(e2e): ban unit tests under tests/e2e by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33852](https://github.com/BerriAI/litellm/pull/33852)
- fix(router): preserve Azure Entra ID params in reusable credentials by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40889](https://github.com/BerriAI/litellm/pull/40889)
- docs(user endpoints): remove unsupported soft\_budget param from user docstrings by [@​shivamrawat1](https://github.com/shivamrawat1) in [#​36585](https://github.com/BerriAI/litellm/pull/36585)
- feat(friendli): auto-sync Friendli model metadata into price registry by [@​Lee-Si-Yoon](https://github.com/Lee-Si-Yoon) in [#​35918](https://github.com/BerriAI/litellm/pull/35918)
- build(deps): bump smol-toml to 1.8.0 to clear GHSA-7w5x-hrqm-74c2 in osv-scan by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40478](https://github.com/BerriAI/litellm/pull/40478)
- fix(bedrock\_mantle): price GovCloud regions from the regional cost row and accept region-prefixed model names by [@​mateo-berri](https://github.com/mateo-berri) in [#​39846](https://github.com/BerriAI/litellm/pull/39846)
- chore(ci): remerge internal staging by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​40943](https://github.com/BerriAI/litellm/pull/40943)
- test(auth): freeze the cache clock in auth prefetch tests by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40996](https://github.com/BerriAI/litellm/pull/40996)
- feat(jwt): allow virtual\_key\_claim\_field per issuer by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40927](https://github.com/BerriAI/litellm/pull/40927)
- fix(cost): bill cached realtime audio tokens at the audio cache-read rate by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40627](https://github.com/BerriAI/litellm/pull/40627)
- perf(logging): skip correlation contextvar stamping when request\_correlation\_in\_logs is off by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41054](https://github.com/BerriAI/litellm/pull/41054)
- feat(pricing): add azure gpt-chat-latest rates and drop retired friendliai llama-3.1 entries by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40976](https://github.com/BerriAI/litellm/pull/40976)
- build(deps): re-suppress GHSA-h7x2-h6g9-p789 in osv-scan on main, mlflow still has no fixed release by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41104](https://github.com/BerriAI/litellm/pull/41104)
- fix(otel): cap per-index OpenInference message attributes span-wide by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40562](https://github.com/BerriAI/litellm/pull/40562)
- feat(proxy): add general\_settings.allowed\_file\_extensions for /v1/files uploads by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41106](https://github.com/BerriAI/litellm/pull/41106)
- fix(proxy): forward provider request id headers on mapped error responses by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40925](https://github.com/BerriAI/litellm/pull/40925)
- fix(router): name the all-deployments-in-cooldown error on 429 responses by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40995](https://github.com/BerriAI/litellm/pull/40995)
- fix(ui): show the team alias on the model info page and in its raw JSON by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40992](https://github.com/BerriAI/litellm/pull/40992)
- feat(proxy): honor LITELLM\_DISABLE\_ACCESS\_LOG\_PATHS to drop noisy uvicorn access log lines by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41096](https://github.com/BerriAI/litellm/pull/41096)
- fix(prometheus): label pre-call rate limit failures with the resolved api\_provider by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41059](https://github.com/BerriAI/litellm/pull/41059)
- perf(proxy): serialize /model/info listing once with orjson by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41114](https://github.com/BerriAI/litellm/pull/41114)
- fix(utils): stop wrapper\_async submitting the sync success handler twice by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41115](https://github.com/BerriAI/litellm/pull/41115)
- fix(redis): log a timeout streak once per interval instead of one line per cache call by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40817](https://github.com/BerriAI/litellm/pull/40817)
- fix(router): record flat retry attempts and cap retries from attempted\_retries by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40930](https://github.com/BerriAI/litellm/pull/40930)
- refactor(prometheus): source PROXY\_LLM\_PROVIDER\_FALLBACK from litellm.constants by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41118](https://github.com/BerriAI/litellm/pull/41118)
- fix(proxy): hide default credentials login hint when UI\_PASSWORD is set by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41107](https://github.com/BerriAI/litellm/pull/41107)
- fix(cli): show routed models and session stats for LLM API keys by [@​tin-berri](https://github.com/tin-berri) in [#​41116](https://github.com/BerriAI/litellm/pull/41116)
- fix(bedrock/realtime): propagate deferred Nova Sonic stream failures to the router by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41064](https://github.com/BerriAI/litellm/pull/41064)
- fix(proxy): keep org admins' own team memberships in other orgs visible on team list by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41086](https://github.com/BerriAI/litellm/pull/41086)
- feat(model\_info): provider-scoped fill\_missing\_for\_providers backfill from fallback generalization rules by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41093](https://github.com/BerriAI/litellm/pull/41093)
- fix(auth): load team membership once per request and skip prisma on an L1 hit by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41102](https://github.com/BerriAI/litellm/pull/41102)
- refactor(harness): expand independent trace coverage by [@​yujonglee-berri](https://github.com/yujonglee-berri) in [#​41120](https://github.com/BerriAI/litellm/pull/41120)
- fix(proxy): release max\_parallel\_requests slot when a realtime session ends without LLM callbacks by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41113](https://github.com/BerriAI/litellm/pull/41113)
- fix(router): cool down team deployments on 429 when a sibling serves the same public model by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40991](https://github.com/BerriAI/litellm/pull/40991)
- fix(ui): move tags typed into key metadata JSON into the Tags field by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41023](https://github.com/BerriAI/litellm/pull/41023)
- fix(ui): let team admins grant a team all proxy models by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​40196](https://github.com/BerriAI/litellm/pull/40196)
- chore(lint): graduate 12 rules from the strict-gate ratchet by [@​HUAHAODIA](https://github.com/HUAHAODIA) in [#​41048](https://github.com/BerriAI/litellm/pull/41048)
- test: add dedicated CircleCI integration contract foundation by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41066](https://github.com/BerriAI/litellm/pull/41066)
- fix(utils): keep litellm params out of provider request bodies by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41018](https://github.com/BerriAI/litellm/pull/41018)
- fix(openai): keep extra\_headers out of the chat request body on the httpx handler path by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41141](https://github.com/BerriAI/litellm/pull/41141)
- test: cover persisted updates and warmed authorization policies by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41070](https://github.com/BerriAI/litellm/pull/41070)
- fix(proxy): resolve x-litellm-call-id from response metadata when routes omit call\_id by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41056](https://github.com/BerriAI/litellm/pull/41056)
- chore(prices): sync Vertex AI prices: 14 models by [@​berriai-litellm-provider-info-sync](https://github.com/berriai-litellm-provider-info-sync)\[bot] in [#​40955](https://github.com/BerriAI/litellm/pull/40955)
- ci(codeql): exclude noisy Python quality queries by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41142](https://github.com/BerriAI/litellm/pull/41142)
- feat(proxy): predict prompt-cache costs across deployments by [@​tin-berri](https://github.com/tin-berri) in [#​40877](https://github.com/BerriAI/litellm/pull/40877)
- fix(prompt\_security): keep polling file sanitization through non-terminal statuses by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41131](https://github.com/BerriAI/litellm/pull/41131)
- fix(health): skip background health check DB writes when the latest-row read fails by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41145](https://github.com/BerriAI/litellm/pull/41145)
- feat(model\_armor): logging\_only mode scans completed streams after delivery by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40702](https://github.com/BerriAI/litellm/pull/40702)
- fix(bedrock guardrails): derive contextual grounding source and query from plain messages by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41132](https://github.com/BerriAI/litellm/pull/41132)
- fix(cli): drop enum.StrEnum so the CLI imports on Python 3.10 by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41046](https://github.com/BerriAI/litellm/pull/41046)
- fix(responses): route mid-stream error events through exception\_type so content\_policy\_fallbacks fire by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40988](https://github.com/BerriAI/litellm/pull/40988)
- fix(cost): bill gemini-embedding-2 per token and stop double charging audio by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41157](https://github.com/BerriAI/litellm/pull/41157)
- test: bind management E2E callers and isolate JWT actors by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​40892](https://github.com/BerriAI/litellm/pull/40892)
- fix(headroom): protect cache\_control-marked rows anywhere in history by [@​rad-p44](https://github.com/rad-p44) in [#​40315](https://github.com/BerriAI/litellm/pull/40315)
- test: add strict stateless provider replay identity by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41149](https://github.com/BerriAI/litellm/pull/41149)
- fix(ci): test checked-out model pricing in unit jobs by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41181](https://github.com/BerriAI/litellm/pull/41181)
- fix(guardrails): write per-message guardrail rewrites back onto Responses input items by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40939](https://github.com/BerriAI/litellm/pull/40939)
- fix(proxy): log the provider usage on deferred /v1/messages calls and price cache writes without a creation rate by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41172](https://github.com/BerriAI/litellm/pull/41172)
- fix(guardrails): record not\_run evaluation when scoping leaves nothing to scan by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​39050](https://github.com/BerriAI/litellm/pull/39050)
- fix(responses): preserve provider affinity by [@​AaronHowell](https://github.com/AaronHowell) in [#​40228](https://github.com/BerriAI/litellm/pull/40228)
- fix(sdk): keep body and proxy headers on BadRequestError mapped from a litellm\_proxy 400 by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40994](https://github.com/BerriAI/litellm/pull/40994)
- fix(responses): hoist Codex additional\_tools input items into the chat bridge tools by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40989](https://github.com/BerriAI/litellm/pull/40989)
- fix(router): honor team and key provider weights by [@​tin-berri](https://github.com/tin-berri) in [#​41072](https://github.com/BerriAI/litellm/pull/41072)
- test(e2e): verify streamed answers and tool continuation by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41194](https://github.com/BerriAI/litellm/pull/41194)
- fix(cli): label router costs and simplify the routed-model header by [@​tin-berri](https://github.com/tin-berri) in [#​41186](https://github.com/BerriAI/litellm/pull/41186)
- test(spend): reconcile concurrent requests and daily activity by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41188](https://github.com/BerriAI/litellm/pull/41188)
- fix(guardrails): scan the Anthropic top-level system prompt and tool\_use arguments by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40984](https://github.com/BerriAI/litellm/pull/40984)
- fix(router): count num\_retries\_per\_request across fallback hops by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41191](https://github.com/BerriAI/litellm/pull/41191)
- fix(bedrock): grant rerank, retrieve, agent, and agentcore actions in the web identity session policy by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41168](https://github.com/BerriAI/litellm/pull/41168)
- fix(vertex-live): bill Gemini Live sessions end to end (internal copy of [#​37075](https://github.com/BerriAI/litellm/issues/37075)) by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40915](https://github.com/BerriAI/litellm/pull/40915)
- fix(health): resolve litellm\_credential\_name in realtime health checks by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41173](https://github.com/BerriAI/litellm/pull/41173)
- feat(proxy): unified custom\_key\_policy hook for key generate, update and regenerate by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40921](https://github.com/BerriAI/litellm/pull/40921)
- fix(proxy): enforce custom\_key\_update policy on /key/regenerate by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40695](https://github.com/BerriAI/litellm/pull/40695)
- fix(router): preserve session model choice within each complexity tier by [@​tin-berri](https://github.com/tin-berri) in [#​41174](https://github.com/BerriAI/litellm/pull/41174)
- test(pricing): let synced GovCloud Bedrock rows cite the AWS price list by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41263](https://github.com/BerriAI/litellm/pull/41263)
- docs(github): ask for interactive coding-tool proof in the PR template by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41257](https://github.com/BerriAI/litellm/pull/41257)
- feat(proxy): add POST /management/v1/users/bulk for batched user and team membership creation by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41028](https://github.com/BerriAI/litellm/pull/41028)
- fix(credentials): answer 409 on a credential name collision, make Terraform adoption opt-in by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​40917](https://github.com/BerriAI/litellm/pull/40917)
- feat(proxy): add POST /management/v1/users/bulk\_delete and POST /management/v1/teams/{team\_id}/members/bulk\_delete by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41039](https://github.com/BerriAI/litellm/pull/41039)
- fix(proxy): list directly assigned team models in model access errors by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41256](https://github.com/BerriAI/litellm/pull/41256)
- feat(auto-router): allow opted-in team members to manage their routers by [@​tin-berri](https://github.com/tin-berri) in [#​41175](https://github.com/BerriAI/litellm/pull/41175)
- build(rust-bridge): add typed \_native stub and validate it with mypy.stubtest by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41180](https://github.com/BerriAI/litellm/pull/41180)
- feat(guardrails): add new upstream presidio pii entities including german set by [@​MvdB](https://github.com/MvdB) in [#​36775](https://github.com/BerriAI/litellm/pull/36775)
- fix(responses): filter bridged kwargs like the native Responses path by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41144](https://github.com/BerriAI/litellm/pull/41144)
- test(e2e): cover the reliability retry, cooldown, fallback, and routing-strategy cells by [@​mateo-berri](https://github.com/mateo-berri) in [#​39857](https://github.com/BerriAI/litellm/pull/39857)
- fix(anthropic): add the per-turn-control beta when a message carries output\_config by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41189](https://github.com/BerriAI/litellm/pull/41189)
- fix(router): bind per-request routing\_strategy override selectors to the request's callbacks by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41178](https://github.com/BerriAI/litellm/pull/41178)
- feat(proxy): bind JWT claims to registered agents via agent\_id\_jwt\_field by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40904](https://github.com/BerriAI/litellm/pull/40904)
- fix(proxy): enforce organization budgets when max\_budget is 0 by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​41271](https://github.com/BerriAI/litellm/pull/41271)
- fix(alerting): send llm\_exceptions Slack alert for 5xx HTTPException and ProxyException by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41125](https://github.com/BerriAI/litellm/pull/41125)
- fix(headroom): protect the cached prefix through the last cache\_control breakpoint by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41161](https://github.com/BerriAI/litellm/pull/41161)
- fix(utils): cache custom HuggingFace tokenizers across /utils/token\_counter requests by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41216](https://github.com/BerriAI/litellm/pull/41216)
- fix(router): keep weighted routing when a deployment id equals a model\_name by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41156](https://github.com/BerriAI/litellm/pull/41156)
- feat(router): add capability classifier as Fuse foundation by [@​tin-berri](https://github.com/tin-berri) in [#​41270](https://github.com/BerriAI/litellm/pull/41270)
- fix(proxy): keep access-group raw SQL writes on the writer while writer\_unavailable is stale by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41283](https://github.com/BerriAI/litellm/pull/41283)
- fix(prometheus): count 401 auth failures in litellm\_proxy\_failed\_requests\_metric by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41170](https://github.com/BerriAI/litellm/pull/41170)
- test: drop tests that pin vendor facts and add the CLAUDE.md rule by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41269](https://github.com/BerriAI/litellm/pull/41269)
- fix(proxy): run the remaining inline token counts off the event loop by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40262](https://github.com/BerriAI/litellm/pull/40262)
- fix(proxy): log blocked streaming guardrail responses as failures, not success by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40191](https://github.com/BerriAI/litellm/pull/40191)
- feat(proxy): add tpd\_limit (tokens per day) for batch submissions by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40997](https://github.com/BerriAI/litellm/pull/40997)
- fix(proxy): reconcile budget reservation before enqueuing spend to the DB by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40310](https://github.com/BerriAI/litellm/pull/40310)
- fix(xai): stop sending web\_search\_options to xAI's retired Live Search path by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​38278](https://github.com/BerriAI/litellm/pull/38278)
- feat(terraform): add tpm\_limit, rpm\_limit, budget\_duration, allowed\_models to litellm\_team\_member\_add by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​38682](https://github.com/BerriAI/litellm/pull/38682)
- fix(rerank): bill Vertex search\_units from input records and give every rerank response a unique id by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35180](https://github.com/BerriAI/litellm/pull/35180)
- fix(router): stop counting caller-set timeout 408s toward deployment cooldown by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41230](https://github.com/BerriAI/litellm/pull/41230)
- feat(router): add Fuse V2 classifier after capability forecasting by [@​tin-berri](https://github.com/tin-berri) in [#​41272](https://github.com/BerriAI/litellm/pull/41272)
- fix(proxy): keep client User-Agent on auth failure spend logs by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41291](https://github.com/BerriAI/litellm/pull/41291)
- fix(proxy): reset budgets by decrementing pre-reset spend instead of zeroing rows by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41279](https://github.com/BerriAI/litellm/pull/41279)
- fix(xai): honor nested web\_search filters on the xAI Responses API by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​38268](https://github.com/BerriAI/litellm/pull/38268)
- fix(router): stop registering a caller-supplied credential as a router deployment by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​41289](https://github.com/BerriAI/litellm/pull/41289)
- fix(router): accept custom\_provider\_map providers before the first completion call by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41300](https://github.com/BerriAI/litellm/pull/41300)
- fix(proxy): return 400 instead of 500 for lone surrogate escapes in request body by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41297](https://github.com/BerriAI/litellm/pull/41297)
- fix(langsmith): keep events appended during an in-flight flush instead of clearing them by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41288](https://github.com/BerriAI/litellm/pull/41288)
- fix(logging): track spend for streams a deployment hook converted to non-streaming by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41171](https://github.com/BerriAI/litellm/pull/41171)
- fix(bedrock): sanitize client tool\_call ids to Bedrock toolUseId constraints by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40872](https://github.com/BerriAI/litellm/pull/40872)
- fix(passthrough): attribute Vertex passthrough successes to the resolved router deployment by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41307](https://github.com/BerriAI/litellm/pull/41307)
- feat(ui): persist Models table search, filters, sort and page in the URL by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41296](https://github.com/BerriAI/litellm/pull/41296)
- fix(jwt-auth): scope JWT key mappings by issuer to prevent cross-issuer collisions by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​41281](https://github.com/BerriAI/litellm/pull/41281)
- feat(openai): add openai\_system\_messages\_first to put system messages first for prompt caching by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41304](https://github.com/BerriAI/litellm/pull/41304)
- feat(ui): add custom request headers to the API Playground by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41309](https://github.com/BerriAI/litellm/pull/41309)
- feat(cli): sync Codex /model picker from proxy /v1/models in lite codex by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40476](https://github.com/BerriAI/litellm/pull/40476)
- chore: bump litellm-enterprise 0.1.67 -> 0.1.68, litellm-proxy-extras 0.4.97 -> 0.4.98, litellm 1.102.0 -> 1.103.0 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41321](https://github.com/BerriAI/litellm/pull/41321)
- feat: add aihubmix provider pricing entries by [@​IToSSc](https://github.com/IToSSc) in [#​41179](https://github.com/BerriAI/litellm/pull/41179)
- feat(auto-router): add per-model Fast mode toggle by [@​tin-berri](https://github.com/tin-berri) in [#​41282](https://github.com/BerriAI/litellm/pull/41282)
- fix(proxy): include litellm\_call\_id in LLM API exception logs by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41205](https://github.com/BerriAI/litellm/pull/41205)
- fix(proxy): keep yaml pass-through endpoints visible to auth after db overlay by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41303](https://github.com/BerriAI/litellm/pull/41303)
- fix(proxy): resolve router\_settings.model\_group\_alias before key/team model auth by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41308](https://github.com/BerriAI/litellm/pull/41308)
- fix(ui): block usage export and flag the range when a spend page fails by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​41294](https://github.com/BerriAI/litellm/pull/41294)
- fix(vertex\_ai): bill Gemini Omni Interactions usage and Veo sampleCount on passthrough by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41322](https://github.com/BerriAI/litellm/pull/41322)
- fix(proxy): honor LITELLM\_LOG for uvicorn and proxy extras loggers by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41306](https://github.com/BerriAI/litellm/pull/41306)
- fix(cost): price native Responses WebSocket turns at their returned service\_tier by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41318](https://github.com/BerriAI/litellm/pull/41318)
- fix(proxy): key model rpm/tpm override takes precedence over team model limit by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41302](https://github.com/BerriAI/litellm/pull/41302)
- fix(proxy): track per-member organization spend by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41255](https://github.com/BerriAI/litellm/pull/41255)
- feat(proxy): add /nvidia\_nim passthrough route for NIM object detection and OCR /v1/infer by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41316](https://github.com/BerriAI/litellm/pull/41316)
- feat(model\_info): add provider-neutral Gemini 2.5+ chat baseline fallback generalization by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41320](https://github.com/BerriAI/litellm/pull/41320)
- fix(spend): sum multi-round session duration in logs UI by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35388](https://github.com/BerriAI/litellm/pull/35388)
- feat(router): limit unlicensed Capability and Fuse v2 routers to one each by [@​tin-berri](https://github.com/tin-berri) in [#​41326](https://github.com/BerriAI/litellm/pull/41326)
- fix(guardrails): resolve caller identity from metadata buckets in custom code guardrail by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41126](https://github.com/BerriAI/litellm/pull/41126)
- fix(e2e): onboard dashboard users through invitations by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41319](https://github.com/BerriAI/litellm/pull/41319)
- feat(guardrails): add Microsoft Agent 365 MCP tool-call guardrail by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​38241](https://github.com/BerriAI/litellm/pull/38241)
- test: drop remaining tests that pin cost-map vendor facts by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41298](https://github.com/BerriAI/litellm/pull/41298)
- feat(ui): show average response time per model in usage model activity by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41313](https://github.com/BerriAI/litellm/pull/41313)
- fix(proxy): preserve Anthropic pricing modifiers in router savings by [@​tin-berri](https://github.com/tin-berri) in [#​41341](https://github.com/BerriAI/litellm/pull/41341)
- feat(guardrails): support pre\_call and during\_call modes for llm\_as\_a\_judge by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41128](https://github.com/BerriAI/litellm/pull/41128)
- fix(gemini): propagate the provider's modelVersion to the response model by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41338](https://github.com/BerriAI/litellm/pull/41338)
- fix(fireworks-ai): bill cache-write, reasoning and audio tokens via the shared cost calculator by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41339](https://github.com/BerriAI/litellm/pull/41339)
- feat(guardrails): singulr v2 API contract with logging\_only, pre\_mcp\_call and post\_mcp\_call by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​41329](https://github.com/BerriAI/litellm/pull/41329)
- ci(image-scan): ignore zlib CVE-2026-85091 until Wolfi ships the fix by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41353](https://github.com/BerriAI/litellm/pull/41353)
- feat(e2e): reuse exact provider responses for 24 hours by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41346](https://github.com/BerriAI/litellm/pull/41346)
- fix(xai): keep 'instructions' on the xAI Responses API so system messages survive web search by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​38254](https://github.com/BerriAI/litellm/pull/38254)
- feat(ui): configure capability and Fuse v2 classifiers by [@​tin-berri](https://github.com/tin-berri) in [#​41315](https://github.com/BerriAI/litellm/pull/41315)
- fix(anthropic): tolerate message\_delta events without usage when streaming by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41336](https://github.com/BerriAI/litellm/pull/41336)
- test(router): ignore deployment-selection logs in the fallback log assertion by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41358](https://github.com/BerriAI/litellm/pull/41358)
- test(proxy): assert budget resets decrement the cleared spend by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41359](https://github.com/BerriAI/litellm/pull/41359)
- fix(e2e): expect models filters to persist after reload by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41348](https://github.com/BerriAI/litellm/pull/41348)
- fix(e2e): record cookie-setting provider responses and keep prompt-caching tests live by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41366](https://github.com/BerriAI/litellm/pull/41366)
- fix(responses): recount tokens when a streamed response completes without usage by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41337](https://github.com/BerriAI/litellm/pull/41337)
- fix(ui): simplify Capability and Fuse advanced routing options by [@​tin-berri](https://github.com/tin-berri) in [#​41371](https://github.com/BerriAI/litellm/pull/41371)
- fix(mcp): authorize JWT OAuth credential persistence by [@​joshua-berri](https://github.com/joshua-berri) in [#​41314](https://github.com/BerriAI/litellm/pull/41314)
- feat(router): stream shadow traffic and fan out silent\_model to multiple targets by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41368](https://github.com/BerriAI/litellm/pull/41368)
- perf(content\_filter): scan a bounded window per streamed chunk by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41407](https://github.com/BerriAI/litellm/pull/41407)
- fix(proxy): hide model allowlist from client-facing model access denied errors by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41310](https://github.com/BerriAI/litellm/pull/41310)
- feat(http): opt-in outbound HTTP/2 for httpx clients by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41268](https://github.com/BerriAI/litellm/pull/41268)
- refactor(rust): remove gateway, config, router, realtime, and Rust trace-parity instrumentation by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41432](https://github.com/BerriAI/litellm/pull/41432)
- fix(guardrails): don't add post\_call output scan for MCP-only Presidio modes by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40571](https://github.com/BerriAI/litellm/pull/40571)
- chore(prices): sync Azure, Azure AI, Gemini, OpenAI, Bedrock, Together AI, Fireworks and Vertex prices: 278 models, 59 new, 30 deprecated by [@​berriai-litellm-provider-info-sync](https://github.com/berriai-litellm-provider-info-sync)\[bot] in [#​41154](https://github.com/BerriAI/litellm/pull/41154)
- fix(rag): forward retrieval\_filter from retrieval\_config to vector store search by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34427](https://github.com/BerriAI/litellm/pull/34427)
- refactor(rust): extract auth and cache crates by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41464](https://github.com/BerriAI/litellm/pull/41464)
- fix(proxy): default litellm\_trace\_id to the OTel server span trace id by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41386](https://github.com/BerriAI/litellm/pull/41386)
- chore(codeowners): add ryan and kerry as owners of the cost map by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41333](https://github.com/BerriAI/litellm/pull/41333)
- fix(responses): guard empty-choices chunks in the Responses API streaming bridge by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34455](https://github.com/BerriAI/litellm/pull/34455)
- chore(prices): sync Google Gemini prices: 22 models by [@​berriai-litellm-provider-info-sync](https://github.com/berriai-litellm-provider-info-sync)\[bot] in [#​41457](https://github.com/BerriAI/litellm/pull/41457)
- fix(bedrock): forward userContext in Knowledge Base Retrieve requests by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41475](https://github.com/BerriAI/litellm/pull/41475)
- ci(rust): split rust jobs, use nextest and Swatinem/rust-cache by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41480](https://github.com/BerriAI/litellm/pull/41480)
- fix(fireworks\_ai): flatten dict-form reasoning\_effort to its effort string by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41335](https://github.com/BerriAI/litellm/pull/41335)
- fix(proxy): never forward the LiteLLM virtual key to Anthropic on the /anthropic passthrough by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41340](https://github.com/BerriAI/litellm/pull/41340)
- fix(proxy): rename AWS Secrets Manager secret when key alias changes by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41468](https://github.com/BerriAI/litellm/pull/41468)
- feat(otel): promote nested request metadata keys to litellm.metadata.\* span attributes by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41462](https://github.com/BerriAI/litellm/pull/41462)
- fix(proxy): sync AWS Secrets Manager on body-less key regenerate and key alias changes by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41458](https://github.com/BerriAI/litellm/pull/41458)
- fix(http\_handler): keep a handler alive while a response it issued is still reading by [@​max-sixty](https://github.com/max-sixty) in [#​34829](https://github.com/BerriAI/litellm/pull/34829)
- fix(bedrock): make prompt caching work on the Nova InvokeModel route by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41343](https://github.com/BerriAI/litellm/pull/41343)
- ci(migrations): flag defaulted ADD COLUMN on request-log tables by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41460](https://github.com/BerriAI/litellm/pull/41460)
- feat(prometheus): add customer (end\_user) budget gauges by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41472](https://github.com/BerriAI/litellm/pull/41472)
- fix(otel): drop None metric and event attributes before OTLP export by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36815](https://github.com/BerriAI/litellm/pull/36815)
- fix(anthropic): carry the served model from message\_start onto stream chunks by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41446](https://github.com/BerriAI/litellm/pull/41446)
- fix(models): rolling registry audit: Gemini latest aliases, Nova cache pricing, OpenRouter/Together sync, Mistral GLM 5.3, Azure snapshots, Grok caching by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41112](https://github.com/BerriAI/litellm/pull/41112)
- fix(router): count TPM/RPM usage before building rate-limit headers by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41474](https://github.com/BerriAI/litellm/pull/41474)
- feat(guardrails): release buffered stream chunks after each passing scan by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41425](https://github.com/BerriAI/litellm/pull/41425)
- fix!: re-check budget on router fallback targets by [@​runjivu](https://github.com/runjivu) in [#​41379](https://github.com/BerriAI/litellm/pull/41379)
- refactor(ocr): move file preparation from the python bridge into litellm-core by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41489](https://github.com/BerriAI/litellm/pull/41489)
- feat(s3): add s3\_log\_prompts\_only option to log prompts without responses by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41327](https://github.com/BerriAI/litellm/pull/41327)
- feat(team): team-level model\_max\_budget with key-level overrides by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41330](https://github.com/BerriAI/litellm/pull/41330)
- feat(keys): filter /key/list by active, expired, revoked or deleted status and serve deleted keys from /key/info by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41311](https://github.com/BerriAI/litellm/pull/41311)
- feat(proxy): expose lifetime total\_spend on virtual keys by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41403](https://github.com/BerriAI/litellm/pull/41403)
- fix(proxy): release completed max-parallel slots promptly by [@​elifozdamar](https://github.com/elifozdamar) in [#​40843](https://github.com/BerriAI/litellm/pull/40843)
- feat(ui): accept ssh clone urls when registering a skill by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35418](https://github.com/BerriAI/litellm/pull/35418)
- fix(prices): dedupe Nova cache\_read\_input\_token\_cost keys left by a text merge by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41496](https://github.com/BerriAI/litellm/pull/41496)
- fix(otel): propagate W3C trace context on HTTP and WebSocket passthrough by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40669](https://github.com/BerriAI/litellm/pull/40669)
- fix(proxy): remove duplicate user budget hook that 429'd zero-cost models by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41345](https://github.com/BerriAI/litellm/pull/41345)
- test(logging): pick this test's own records out of the shared log batch by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41487](https://github.com/BerriAI/litellm/pull/41487)
- test(together\_ai): move request-shape checks to the mapped file, drop the live ones by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41360](https://github.com/BerriAI/litellm/pull/41360)
- feat(ui): shared URL-state layer for tables and tabs by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​41331](https://github.com/BerriAI/litellm/pull/41331)
- feat(e2e): make the provider cache reusable across builds and mount Bedrock behind it by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41402](https://github.com/BerriAI/litellm/pull/41402)
- fix(mcp): fail closed on missing upstream credentials by [@​joshua-berri](https://github.com/joshua-berri) in [#​41364](https://github.com/BerriAI/litellm/pull/41364)
- feat(rust): scaffold Redis cache crate by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41501](https://github.com/BerriAI/litellm/pull/41501)
- fix(dashscope): forward reasoning\_effort to the provider by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​37506](https://github.com/BerriAI/litellm/pull/37506)
- fix(proxy): carry litellm\_call\_id through endpoint specific error logs and failure responses by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41356](https://github.com/BerriAI/litellm/pull/41356)
- fix(proxy): retry rate-limit fallbacks from a pristine request snapshot by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40596](https://github.com/BerriAI/litellm/pull/40596)
- fix(gemini): map minimal thinking to low for Gemini 3.7 and 3.8 Flash by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41201](https://github.com/BerriAI/litellm/pull/41201)
- fix(proxy): stop forwarding LiteLLM credential headers on Bedrock agent-runtime passthrough by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41504](https://github.com/BerriAI/litellm/pull/41504)
- fix(streaming): estimate interrupted Anthropic stream usage from reasoning\_content by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41503](https://github.com/BerriAI/litellm/pull/41503)
- fix(azure\_ai): route Responses API to native /openai/v1/responses for Foundry Models by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33856](https://github.com/BerriAI/litellm/pull/33856)
- fix(proxy): show all model groups to proxy admins in /model\_group/info by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41094](https://github.com/BerriAI/litellm/pull/41094)
- feat(proxy): let proxy admins choose which team fields team admins may edit by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​39996](https://github.com/BerriAI/litellm/pull/39996)
- fix(bedrock): neutralize orphaned tool blocks instead of raising or injecting a dummy tool (internal copy of [#​31400](https://github.com/BerriAI/litellm/issues/31400)) by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41513](https://github.com/BerriAI/litellm/pull/41513)
- feat(ui): persist organizations and projects list, detail tab and key table state in the URL by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​41445](https://github.com/BerriAI/litellm/pull/41445)
- fix(bedrock\_mantle): accept and forward verbosity on gpt-5.x chat completions by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41509](https://github.com/BerriAI/litellm/pull/41509)
- test: cover database transactions and persisted accounting by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41073](https://github.com/BerriAI/litellm/pull/41073)
- test: provider wire contracts, streaming and recovery by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41075](https://github.com/BerriAI/litellm/pull/41075)
- fix(mcp): count admin static headers as api\_key credential slots by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41514](https://github.com/BerriAI/litellm/pull/41514)
- ci: auto-merge provider-info-sync PRs when CI, Greptile and Bugbot are clean by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41494](https://github.com/BerriAI/litellm/pull/41494)
- feat(rust): add standalone framing crate by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41500](https://github.com/BerriAI/litellm/pull/41500)
- fix(proxy): enforce tag budgets for tags added by guardrails by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40842](https://github.com/BerriAI/litellm/pull/40842)
- fix(utils): run post-call deployment hook on converted chat streams by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41495](https://github.com/BerriAI/litellm/pull/41495)
- fix(e2e): bind provider-cache recordings to the deployment's test, not the serving process by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41520](https://github.com/BerriAI/litellm/pull/41520)
- test: add extension and browser integration contracts by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41078](https://github.com/BerriAI/litellm/pull/41078)
- fix(logging): scan each log record once and collapse base64 payloads before the secret regex by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40934](https://github.com/BerriAI/litellm/pull/40934)
- fix(spend\_tracking): attribute router-rejected requests to the model group provider by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41507](https://github.com/BerriAI/litellm/pull/41507)
- feat(router): discover token limits for hosted OpenAI-compatible models by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41508](https://github.com/BerriAI/litellm/pull/41508)
- feat(proxy): let team admins edit rpm\_limit and max\_budget when enabled by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​41525](https://github.com/BerriAI/litellm/pull/41525)
- fix(otel): fit per-index OpenInference messages to the span's remaining attribute budget by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41498](https://github.com/BerriAI/litellm/pull/41498)
- test(aws): verify rotated secret value by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41524](https://github.com/BerriAI/litellm/pull/41524)
- test(e2e): read a deleted key back as deleted, not as a 404 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41551](https://github.com/BerriAI/litellm/pull/41551)
- fix(otel v2): map the caller's Langfuse user, session and tags onto the root and generation spans by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41140](https://github.com/BerriAI/litellm/pull/41140)
- test: fix seven tests left stale by [#​41311](https://github.com/BerriAI/litellm/issues/41311), [#​41337](https://github.com/BerriAI/litellm/issues/41337), [#​39996](https://github.com/BerriAI/litellm/issues/39996), [#​41310](https://github.com/BerriAI/litellm/issues/41310), [#​41289](https://github.com/BerriAI/litellm/issues/41289) and [#​41315](https://github.com/BerriAI/litellm/issues/41315) by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41527](https://github.com/BerriAI/litellm/pull/41527)
- test(budgets): cover management null handling by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41563](https://github.com/BerriAI/litellm/pull/41563)
- test(e2e): drop the auto-router select "opens below" spec by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41568](https://github.com/BerriAI/litellm/pull/41568)
- fix(guardrails): stream Prompt Security post\_call redactions in incremental\_diff mode by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41558](https://github.com/BerriAI/litellm/pull/41558)
- fix(guardrails): give post-call scans the scoped request conversation and tools by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41220](https://github.com/BerriAI/litellm/pull/41220)
- feat(openrouter): add stealth/union-alpha to the model cost map by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41576](https://github.com/BerriAI/litellm/pull/41576)
- test(management): cover project authorization lifecycle by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41573](https://github.com/BerriAI/litellm/pull/41573)
- feat(rust): map Anthropic Messages transformations by [@​yujonglee-berri](https://github.com/yujonglee-berri) in [#​41531](https://github.com/BerriAI/litellm/pull/41531)
- fix(e2e): clear the three standing errors in the scheduled Buildkite suite by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41616](https://github.com/BerriAI/litellm/pull/41616)
- refactor(rust\_bridge): declarative route catalog and shared runtime selection by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41479](https://github.com/BerriAI/litellm/pull/41479)
- fix(mock\_completion): keep the resolved provider so router custom pricing resolves for azure\_ai deployments by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41623](https://github.com/BerriAI/litellm/pull/41623)
- fix(mcp): restrict health discovery to virtual key grants by [@​joshua-berri](https://github.com/joshua-berri) in [#​41609](https://github.com/BerriAI/litellm/pull/41609)
- fix(mcp): preserve request-selected guardrails during tool execution by [@​joshua-berri](https://github.com/joshua-berri) in [#​41619](https://github.com/BerriAI/litellm/pull/41619)
- refactor(ocr): mirror Python provider layout and preserve tests by [@​yujonglee-berri](https://github.com/yujonglee-berri) in [#​41550](https://github.com/BerriAI/litellm/pull/41550)
- test(fireworks\_ai): stop pinning vision support on minimax-m3 by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41627](https://github.com/BerriAI/litellm/pull/41627)
- perf(spend\_tracking): index LiteLLM\_SpendLogs by (api\_key, startTime) by [@​etiennechabert](https://github.com/etiennechabert) in [#​37983](https://github.com/BerriAI/litellm/pull/37983)
- fix(proxy): reject non-string model with 400 and log its spend as unknown-model by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41633](https://github.com/BerriAI/litellm/pull/41633)
- test(together\_ai): stop pinning successor deprecation status by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41635](https://github.com/BerriAI/litellm/pull/41635)
- chore(prices): sync Together AI prices: 6 models, 6 deprecated \[sync failed: Google Gemini] by [@​berriai-litellm-provider-info-sync](https://github.com/berriai-litellm-provider-info-sync)\[bot] in [#​41570](https://github.com/BerriAI/litellm/pull/41570)
- fix(budgets): page end-user cache invalidation after a budget reset by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​41488](https://github.com/BerriAI/litellm/pull/41488)
- chore: bump litellm-proxy-extras 0.4.98 -> 0.4.99 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41659](https://github.com/BerriAI/litellm/pull/41659)
- fix(tests): resolve the integration support package without run.py's PYTHONPATH by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41373](https://github.com/BerriAI/litellm/pull/41373)
- fix(ui): keep untimed guardrail entries on the request lifecycle by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41374](https://github.com/BerriAI/litellm/pull/41374)
- fix(anthropic-bridge): convert mid-conversation system turns to user turns on /v1/messages to chat completions by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41493](https://github.com/BerriAI/litellm/pull/41493)
- fix(bedrock): support aws-sdk-bedrock-runtime 0.10/0.11 in Bedrock Realtime by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41542](https://github.com/BerriAI/litellm/pull/41542)
- feat(cli): deprecate the litellm-proxy entrypoint in favour of lite by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41673](https://github.com/BerriAI/litellm/pull/41673)
- fix(scim): align pagination `count` validation with RFC 7644 by [@​zachbernstein-sdx](https://github.com/zachbernstein-sdx) in [#​41444](https://github.com/BerriAI/litellm/pull/41444)
- fix(bedrock): never emit Converse cachePoint for OpenAI-family models by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41419](https://github.com/BerriAI/litellm/pull/41419)
- fix(images): stop forwarding the raw image\[] and mask\[] form keys by [@​mateo-berri](https://github.com/mateo-berri) in [#​39512](https://github.com/BerriAI/litellm/pull/39512)
- feat(management\_v1): bulk update team member budgets by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​41632](https://github.com/BerriAI/litellm/pull/41632)
- refactor(rust): extract provider translations by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41690](https://github.com/BerriAI/litellm/pull/41690)
- feat(cli): rename lite autoroute up/down to start/stop, keeping the old names as deprecated aliases by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41672](https://github.com/BerriAI/litellm/pull/41672)
- fix(responses): keep the addressed response id off bridged provider requests by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41689](https://github.com/BerriAI/litellm/pull/41689)
- fix(license): let a wildcard allowed\_features license grant the auto\_router feature by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41684](https://github.com/BerriAI/litellm/pull/41684)
- fix(ui): list every provider in the cache leakage by-model table by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40875](https://github.com/BerriAI/litellm/pull/40875)
- fix(team): keep a forked member budget's reset window and audit bulk member budget writes by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​41686](https://github.com/BerriAI/litellm/pull/41686)
- feat(proxy): add TypeSafe AI Jev evaluate passthrough with registry-priced spend tracking by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41607](https://github.com/BerriAI/litellm/pull/41607)
- test(e2e): cover bedrock batch file upload and create in the us-gov-west-1 partition by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41536](https://github.com/BerriAI/litellm/pull/41536)
- feat(grafana): add all-metrics dashboard and fix stale dashboard\_v2 gauges by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41578](https://github.com/BerriAI/litellm/pull/41578)
- fix(cost): price Azure PTU spillover requests at standard token rates by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41569](https://github.com/BerriAI/litellm/pull/41569)
- build(deps): bump soupsieve to 2.9.2 to clear the osv-scan advisories by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41703](https://github.com/BerriAI/litellm/pull/41703)
- fix(fireworks\_ai): restore supports\_vision on minimax-m3 in the cost map by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41699](https://github.com/BerriAI/litellm/pull/41699)
- feat(policy\_engine): explicit priority for policy attachment execution order by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41571](https://github.com/BerriAI/litellm/pull/41571)
- feat(router): add TypeSafe Jev as a complexity router classifier by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41615](https://github.com…
Note
Separate copy of #39961 by @tin-berri, refreshed against
main. Original commits and authorship are preserved. Fuse V2 follows in #41272, targeting this branchTLDR
Problem this solves:
How it solves it:
User Flow
Before: the developer cannot configure capability-based routing
capabilityconfiguration to POST http://localhost:4000/auto_router/test_routingAfter: the developer can preview and serve capability-based routing
model: auto-capabilityand receive HTTP 200 with the solver answerRelevant issues
Original implementation: #39961
Affected release
Linear ticket
Pre-Submission checklist
Screenshots / Proof of Fix
The before checkout omits
auto-capability, which its configuration schema does not support. The after checkout adds the alias below. Provider deployments, prompts, and endpoint payloads are the sameShared setup and commands
Set
MOONSHOT_API_KEYin the environment and save this asqa/capability.yaml. UseV2_QA_MASTER_KEY=sk-local-classifiers-qafor these localhost-only instancesStart the proxy from each checkout with
PYTHON_DOTENV_DISABLED=1 uv run --no-sync python -m litellm.proxy.proxy_cli --config qa/capability.yaml --host 127.0.0.1 --port PORT. Before uses port 4135; after uses 4136Save these request payloads in
qa/:capability-preview.json{"prompt": "Return only the Python expression that adds two to x", "complexity_router_config": {"classifier_type": "capability", "classifier_llm_config": {"model": "judge", "timeout_ms": 60000, "circuit_breaker_enabled": false}, "tiers": {"SIMPLE": ["efficient"], "REASONING": ["capable"]}, "classification_mode": "user_turn", "adaptive": false, "plan_mode_min_tier": null, "route_housekeeping_to_cheapest_tier": false, "stall_escalation_enabled": false, "enable_context_window_escalation": false, "modality_routing": false, "escalation_keywords": [], "keyword_tier_rules": [], "return_raw_model_name": true, "max_tokens_from_tier_model": false, "capability_classifier_config": {"efficient_tier": "SIMPLE", "capable_tier": "REASONING", "base_threshold": 0.66, "threshold_step": 0.1, "max_output_tokens": 1024, "response_format": "json_schema"}}}capability-chat.json{"model": "auto-capability", "messages": [{"role": "user", "content": "Return only the Python expression that adds two to x"}], "max_tokens": 128}capability-responses.json{"model": "auto-capability", "input": "Return only the Python expression that adds two to x", "max_output_tokens": 128}capability-messages.json{"model": "auto-capability", "messages": [{"role": "user", "content": "Return only the Python expression that adds two to x"}], "max_tokens": 128}Use this curl helper for the steps below:
The after run made paid calls to Moonshot Kimi K3. Solver aliases use thinking disabled/enabled to exercise transport and routing; this run does not measure solver quality. Batches were spaced at least 60 seconds apart for the provider's rate limit
Before (3ac7975)
Routing preview
request http://127.0.0.1:4135 auto_router/test_routing qa/capability-preview.jsonclassifier_typerejected: Input should be 'heuristic', 'heuristic_v2', 'llm', 'custom', 'heuristic_first' or 'hybrid'Chat Completions
request http://127.0.0.1:4135 v1/chat/completions qa/capability-chat.json/chat/completions: Invalid model name passed in model=auto-capability. Call/v1/modelsto view available models for your key.Responses
request http://127.0.0.1:4135 v1/responses qa/capability-responses.json/responses: Invalid model name passed in model=auto-capability. Call/v1/modelsto view available models for your key.Messages
request http://127.0.0.1:4135 v1/messages qa/capability-messages.jsonanthropic_messages: Invalid model name passed in model=auto-capability. Call/v1/modelsto view available models for your key.After (55fc0de)
Routing preview
request http://127.0.0.1:4136 auto_router/test_routing qa/capability-preview.json{"routed_model": "efficient", "cause": "capability_classifier", "classifier_p_solve": 0.98, "classifier_threshold": 0.66}Chat Completions
request http://127.0.0.1:4136 v1/chat/completions qa/capability-chat.jsonchatcmpl-6aa9985eb83a9dd71440356freturnedx + 2frommoonshot/kimi-k3Responses
request http://127.0.0.1:4136 v1/responses qa/capability-responses.jsonresp_5dOeLOtqN03sghXHzTPGNzmGhC1s7JyYqq5ZqhPHFxmxnxZffxMScy1HlZYE0q_doicxjtoNQBsJT_5rRFvqjEAZR8pgJbYA7AJqk39-ZS5NKL3ruZjWpvxx7LVqeMRF4t3qTm_ihOsx2GdkNOS1KN23Jp_3gx8B2E9ugerVyxERTNaEhCR7rDlwAUBRtvmejdxlRXUlDzxNg6uTG3b0kgOf8-qUYKvEMHRvbgmiBshvZqZv6UN2qMFUeJO01fZMSwXVLJEJg4vBv2Y1R6qkpW5BPCANCoCgfBkStpgbHVgbNTDJxcaI6YjbaH20cwmvrRdUbq8eEVS5-nW3KhnqAIY8HXmRQDboqRiIG3F7nN5nYrW0qvtzrBKdv6CdRX2J5s4CoAVoEp000Rkyzb9wIUlJm_cLOsDHERCV1ieCsqvoOwdoy9k=returnedx + 2frommoonshot/kimi-k3Messages
request http://127.0.0.1:4136 v1/messages qa/capability-messages.jsonchatcmpl-6aa998ec45deacfb3b1a0171returnedx + 2frommoonshot/kimi-k3Type
New Feature
Caveats (if any)
Medium
Final Attestation