feat(guardrails): singulr v2 API contract with logging_only, pre_mcp_call and post_mcp_call - #41329
Conversation
|
bugbot run |
Greptile SummaryThis PR updates the Singulr guardrail to its v2 request contract, adds logging-only and MCP lifecycle hooks, uses typed response and MCP payload models, and rejects invalid null verdicts according to
Confidence Score: 5/5The PR appears safe to merge, with no new actionable failures introduced since the previous review Both previous findings were resolved, and the reviewed Singulr files have no changes since the previous review SHA that introduce a new defect Important Files Changed
Reviews (6): Last reviewed commit: "fix(guardrails): run Singulr logging_onl..." | Re-trigger Greptile |
| correlation_id=request_data.get("litellm_call_id"), | ||
| guardrail_scope="response", | ||
| model_name=request_data.get("model"), | ||
| messages=request_data.get("messages"), |
There was a problem hiding this comment.
🟡 Response model identity is lost
When a provider returns a concrete model, _apply_guardrail_on_response sends the request alias instead. process_output_response provides the response model in inputs["model"]. Model-specific Singulr evaluation can use the wrong model.
Learn more
The response translation extracts the provider's returned model and stores it in inputs["model"]. This can differ from the client-facing alias retained in request_data["model"]. The request path already uses inputs["model"], but the response path discards it. Singulr therefore receives inconsistent model identities for the two halves of one call.
Example: A request targets alias support-model, which resolves to gpt-5.4-mini-2026-08-01. The response scan sends support-model although the translation supplied the concrete model. A model-specific Singulr rule or audit record is then associated with the alias instead of the deployed model.
Recommended fix: Prefer inputs.get("model"), matching the request path and the translation contract. Fall back to request_data.get("model") for response handlers that do not populate the input model.
Was this helpful? React with 👍 or 👎 to provide feedback.
There was a problem hiding this comment.
Both scopes deliberately carry the alias the caller requested; the response scope switching to the provider's returned model would make the two halves of one call disagree
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
| async def _apply_guardrail_on_mcp_response( | ||
| self, inputs: GenericGuardrailAPIInputs, texts: Sequence[str], request_data: Mapping[str, Any] | ||
| ) -> GenericGuardrailAPIInputs: | ||
| if not texts: |
There was a problem hiding this comment.
Low: Non-text MCP results bypass the guardrail
An authenticated user can receive image or embedded-resource output from an MCP tool without Singulr inspecting it. The MCP translation only places extracted strings in texts, so this early return skips image-only results and the MCP payload has no field for forwarding those content blocks; include all MCP result content supported by Singulr or fail closed when a result contains content that cannot be scanned.
There was a problem hiding this comment.
The MCP translation hands every guardrail text content only, and Singulr's tool_result contract is text; image results are out of scope here and noted in the PR caveats
There was a problem hiding this comment.
Thanks for the detail. I can't automatically confirm this is safe to dismiss, so I'm leaving the thread open for a maintainer to make the call.
PR overviewThis PR adds the Singulr v2 guardrail API contract, including logging-only behavior and pre- and post-MCP-call processing hooks. One security gap remains in post-call processing: non-text MCP results such as images or embedded resources can bypass Singulr inspection and still be returned to an authenticated user. This is a confirmed guardrail bypass with a limited scope, and no issues have yet been addressed. Open issues (1)
Fixed/addressed: 0 · PR risk: 5/10 |
|
bugbot run |
There was a problem hiding this comment.
Note
This report is out of date. Scroll down for Devin Review's latest report on this PR.
Devin Review found 1 new potential issue.
1 flag not posted on this PR by your GitHub settings — view it in Devin Review. (Configure)
| images: Final = inputs.get("images") | ||
| tools: Final = inputs.get("tools") | ||
|
|
||
| if not messages and not images and not tools: | ||
| verbose_proxy_logger.debug("Singulr: No messages, images, or tools to check after filtering") |
There was a problem hiding this comment.
🟡 Image-only prompts bypass Singulr
For image-only chat prompts, _apply_guardrail_on_request is never reached despite accepting images. process_input_messages dispatches only for text or tool calls. Singulr never evaluates these prompts.
Learn more
The chat translation extracts image URLs into images_to_check, but its dispatch condition ignores that collection. The new Singulr branch can send images and deliberately treats images as sufficient input, yet normal chat requests cannot enter it without text or tool calls. This affects valid multimodal messages whose content contains only image_url blocks.
Example: A user sends one message containing only an image_url. The translation extracts the URL but skips apply_guardrail, while Singulr was expected to receive a request payload containing that image.
Recommended fix: Update process_input_messages to dispatch when images_to_check is non-empty, then add an integration-level regression test through the translation handler rather than calling apply_guardrail directly.
Was this helpful? React with 👍 or 👎 to provide feedback.
There was a problem hiding this comment.
The shared chat translation only dispatches for text or tool calls, for every guardrail; changing that gate is outside this PR and noted in the caveats
|
@greptileai review latest head |
|
bugbot run |
|
bugbot run |
There was a problem hiding this comment.
Note
This report is out of date. Scroll down for Devin Review's latest report on this PR.
Devin Review found 1 new potential issue.
⚠️ 1 issue in files not directly in the diff
⚠️ Embedding inputs bypass logging guardrails
Normalized embedding logs make scratch_input select message dictionaries over raw input. The embeddings translator skips them, so logging-only guardrails inspect nothing
1 flag not posted on this PR by your GitHub settings — view it in Devin Review. (Configure)
|
@greptileai custom_guardrail.py is no longer in this diff and matches litellm_internal_staging, so embeddings under logging_only behave as on staging today, disclosed in Caveats |
|
You're right. |
|
@greptileai review and rescore on the latest head |
|
bugbot run |
…call and post_mcp_call Squash of #37464 (head da298ca) by @aniket-kardile, adopted onto main: v2 gateway payload contract with request, response, mcp_request and mcp_response scopes, typed payload models, proxy user, org and team metadata forwarded to Singulr, and the logging_only, pre_mcp_call and post_mcp_call modes.
…MCP scans off the proxy call type and type the payloads Removes the Singulr async_logging_hook and logging_hook overrides so logging_only runs through CustomGuardrail.async_logging_hook: the response scope reaches Singulr as an assistant message instead of a raw ModelResponse dump, a vendor timeout is recorded as guardrail_failed_to_respond, a request-scope block ends the scan, and the sync success callback thread makes no Singulr call. Decides MCP versus LLM by the proxy logging object's call_type (then the call_type or server-only markers in request_data), never by name, arguments or mcp_tool_name keys a client can put in a chat body. REST /mcp-rest/tools/call pre-scans reach Singulr as mcp_request and a non-mapping arguments value is forwarded as tool_arguments instead of raising. should_block is a strict bool defaulting to false so a null verdict is an invalid response that block_on_error decides; payload fields drop Any for Sequence, Mapping and AssistantMessage types; metadata carries only the keys present; docstrings and section comments removed per the repo comment policy.
050e51e to
1460856
Compare
|
bugbot run |
There was a problem hiding this comment.
Devin Review found 2 new potential issues.
3 flags not posted on this PR by your GitHub settings — view them in Devin Review. (Configure)
There was a problem hiding this comment.
🟡 Malformed verdicts bypass error handling
When Singulr returns a coercible non-boolean, should_block accepts it. The malformed verdict bypasses block_on_error and becomes a block or allow.
(Refers to this code)
Learn more
Pydantic boolean fields are coercive unless strict validation is enabled. The v2 response contract requires a JSON boolean, and _call_api already maps validation failures through block_on_error. Coercible strings and integers never reach that failure path.
Example: A malformed response containing {"should_block": "yes"} validates as True and blocks the request. It was expected to be treated as an invalid response and follow block_on_error.
Recommended fix: Mark should_block strict and add regression cases for coercible strings and integers, not only null.
Was this helpful? React with 👍 or 👎 to provide feedback.
There was a problem hiding this comment.
Singulr emits a JSON boolean or omits the field; every guardrail response model here types verdicts as plain bool, so this stays consistent
| ) | ||
| resolved: Final = ( | ||
| *((field, cls._resolve_metadata_value(request_data=request_data, key=field)) for field in fields), | ||
| ("user_api_key_user_role", cls._resolve_user_role_from_request_data(request_data=request_data)), |
There was a problem hiding this comment.
🟡 User role metadata silently disappears
When unified hooks provide user_api_key_user_role directly, _build_metadata drops it. Singulr then cannot evaluate policies using the caller's role.
Learn more
Unified guardrail dispatch serializes UserAPIKeyAuth into prefixed metadata fields before calling the integration. That produces user_api_key_user_role directly in the metadata bucket, without retaining a nested user_api_key_auth object. The new resolver only handles the nested object shape, which the added tests construct manually.
Example: A caller with role INTERNAL_USER_VIEW_ONLY reaches the unified pre-call hook. The hook stores user_api_key_user_role: "internal_user_view_only", but the Singulr payload omits it.
Recommended fix: Resolve user_api_key_user_role through _resolve_metadata_value first, then retain _resolve_user_role_from_request_data as a fallback for callers that pass the object.
| ("user_api_key_user_role", cls._resolve_user_role_from_request_data(request_data=request_data)), | |
| ( | |
| "user_api_key_user_role", | |
| cls._resolve_metadata_value(request_data=request_data, key="user_api_key_user_role") | |
| or cls._resolve_user_role_from_request_data(request_data=request_data), | |
| ), |
Was this helpful? React with 👍 or 👎 to provide feedback.
There was a problem hiding this comment.
Nothing in litellm or enterprise writes user_api_key_user_role; the proxy sets metadata.user_api_key_auth on every request (litellm_pre_call_utils.py:1628), and 544 of 589 live head payloads carried proxy_admin
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 1460856. Configure here.
…103.0) (#290)
This PR contains the following updates:
| Package | Update | Change |
|---|---|---|
| [ghcr.io/berriai/litellm](https://images.chainguard.dev/directory/image/wolfi-base/overview) ([source](https://github.com/BerriAI/litellm)) | minor | `v1.102.1` → `v1.103.0` |
---
### Release Notes
<details>
<summary>BerriAI/litellm (ghcr.io/berriai/litellm)</summary>
### [`v1.103.0`](https://github.com/BerriAI/litellm/releases/tag/v1.103.0)
[Compare Source](https://github.com/BerriAI/litellm/compare/v1.102.1...v1.103.0)
#### Verify Docker Image Signature
All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0).
**Verify using the pinned commit hash (recommended):**
A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key:
```bash
cosign verify \
--key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \
ghcr.io/berriai/litellm:v1.103.0
```
**Verify using the release tag (convenience):**
Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules:
```bash
cosign verify \
--key https://raw.githubusercontent.com/BerriAI/litellm/v1.103.0/cosign.pub \
ghcr.io/berriai/litellm:v1.103.0
```
Expected output:
```
The following checks were performed on each of these signatures:
- The cosign claims were validated
- The signatures were verified against the specified public key
```
***
#### What's Changed
- fix(responses): translate the reasoning object into a chat-completion reasoning effort by [@​joshgarnett](https://github.com/joshgarnett) in [#​36363](https://github.com/BerriAI/litellm/pull/36363)
- fix(proxy): bound tool and guardrail index create\_many by the spend-log statement budgets by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40561](https://github.com/BerriAI/litellm/pull/40561)
- fix(mcp): require admission for delegated OAuth by [@​joshua-berri](https://github.com/joshua-berri) in [#​40923](https://github.com/BerriAI/litellm/pull/40923)
- fix(logging): log one bounded summary for a burst of timed-out LoggingWorker callbacks by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40912](https://github.com/BerriAI/litellm/pull/40912)
- fix(fireworks): resolve short model names to long cost map keys by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40929](https://github.com/BerriAI/litellm/pull/40929)
- ci: remove main branch source guard by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​40172](https://github.com/BerriAI/litellm/pull/40172)
- chore(ci): promote internal staging to main by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​40942](https://github.com/BerriAI/litellm/pull/40942)
- fix(spend\_logs): store litellm\_call\_id and match it in request\_id lookups by [@​mateo-berri](https://github.com/mateo-berri) in [#​39068](https://github.com/BerriAI/litellm/pull/39068)
- fix(auth): refresh lite login session token grants from the live user and team rows by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​40657](https://github.com/BerriAI/litellm/pull/40657)
- feat(bedrock): support file delete and list for S3-backed managed files by [@​mateo-berri](https://github.com/mateo-berri) in [#​39836](https://github.com/BerriAI/litellm/pull/39836)
- fix(proxy): gate the webhook test alert on proxy admins by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​40814](https://github.com/BerriAI/litellm/pull/40814)
- fix(ui): hide admin write-form tabs on the models page from view-only admins by [@​mateo-berri](https://github.com/mateo-berri) in [#​38867](https://github.com/BerriAI/litellm/pull/38867)
- fix(anthropic-adapter): surface mid-stream provider errors as Anthropic error events by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33352](https://github.com/BerriAI/litellm/pull/33352)
- docs(github): add an Affected release section to the PR template by [@​mateo-berri](https://github.com/mateo-berri) in [#​40618](https://github.com/BerriAI/litellm/pull/40618)
- docs(e2e): ban unit tests under tests/e2e by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33852](https://github.com/BerriAI/litellm/pull/33852)
- fix(router): preserve Azure Entra ID params in reusable credentials by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40889](https://github.com/BerriAI/litellm/pull/40889)
- docs(user endpoints): remove unsupported soft\_budget param from user docstrings by [@​shivamrawat1](https://github.com/shivamrawat1) in [#​36585](https://github.com/BerriAI/litellm/pull/36585)
- feat(friendli): auto-sync Friendli model metadata into price registry by [@​Lee-Si-Yoon](https://github.com/Lee-Si-Yoon) in [#​35918](https://github.com/BerriAI/litellm/pull/35918)
- build(deps): bump smol-toml to 1.8.0 to clear GHSA-7w5x-hrqm-74c2 in osv-scan by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40478](https://github.com/BerriAI/litellm/pull/40478)
- fix(bedrock\_mantle): price GovCloud regions from the regional cost row and accept region-prefixed model names by [@​mateo-berri](https://github.com/mateo-berri) in [#​39846](https://github.com/BerriAI/litellm/pull/39846)
- chore(ci): remerge internal staging by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​40943](https://github.com/BerriAI/litellm/pull/40943)
- test(auth): freeze the cache clock in auth prefetch tests by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40996](https://github.com/BerriAI/litellm/pull/40996)
- feat(jwt): allow virtual\_key\_claim\_field per issuer by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40927](https://github.com/BerriAI/litellm/pull/40927)
- fix(cost): bill cached realtime audio tokens at the audio cache-read rate by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40627](https://github.com/BerriAI/litellm/pull/40627)
- perf(logging): skip correlation contextvar stamping when request\_correlation\_in\_logs is off by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41054](https://github.com/BerriAI/litellm/pull/41054)
- feat(pricing): add azure gpt-chat-latest rates and drop retired friendliai llama-3.1 entries by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40976](https://github.com/BerriAI/litellm/pull/40976)
- build(deps): re-suppress GHSA-h7x2-h6g9-p789 in osv-scan on main, mlflow still has no fixed release by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41104](https://github.com/BerriAI/litellm/pull/41104)
- fix(otel): cap per-index OpenInference message attributes span-wide by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40562](https://github.com/BerriAI/litellm/pull/40562)
- feat(proxy): add general\_settings.allowed\_file\_extensions for /v1/files uploads by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41106](https://github.com/BerriAI/litellm/pull/41106)
- fix(proxy): forward provider request id headers on mapped error responses by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40925](https://github.com/BerriAI/litellm/pull/40925)
- fix(router): name the all-deployments-in-cooldown error on 429 responses by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40995](https://github.com/BerriAI/litellm/pull/40995)
- fix(ui): show the team alias on the model info page and in its raw JSON by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40992](https://github.com/BerriAI/litellm/pull/40992)
- feat(proxy): honor LITELLM\_DISABLE\_ACCESS\_LOG\_PATHS to drop noisy uvicorn access log lines by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41096](https://github.com/BerriAI/litellm/pull/41096)
- fix(prometheus): label pre-call rate limit failures with the resolved api\_provider by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41059](https://github.com/BerriAI/litellm/pull/41059)
- perf(proxy): serialize /model/info listing once with orjson by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41114](https://github.com/BerriAI/litellm/pull/41114)
- fix(utils): stop wrapper\_async submitting the sync success handler twice by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41115](https://github.com/BerriAI/litellm/pull/41115)
- fix(redis): log a timeout streak once per interval instead of one line per cache call by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40817](https://github.com/BerriAI/litellm/pull/40817)
- fix(router): record flat retry attempts and cap retries from attempted\_retries by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40930](https://github.com/BerriAI/litellm/pull/40930)
- refactor(prometheus): source PROXY\_LLM\_PROVIDER\_FALLBACK from litellm.constants by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41118](https://github.com/BerriAI/litellm/pull/41118)
- fix(proxy): hide default credentials login hint when UI\_PASSWORD is set by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41107](https://github.com/BerriAI/litellm/pull/41107)
- fix(cli): show routed models and session stats for LLM API keys by [@​tin-berri](https://github.com/tin-berri) in [#​41116](https://github.com/BerriAI/litellm/pull/41116)
- fix(bedrock/realtime): propagate deferred Nova Sonic stream failures to the router by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41064](https://github.com/BerriAI/litellm/pull/41064)
- fix(proxy): keep org admins' own team memberships in other orgs visible on team list by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41086](https://github.com/BerriAI/litellm/pull/41086)
- feat(model\_info): provider-scoped fill\_missing\_for\_providers backfill from fallback generalization rules by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41093](https://github.com/BerriAI/litellm/pull/41093)
- fix(auth): load team membership once per request and skip prisma on an L1 hit by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41102](https://github.com/BerriAI/litellm/pull/41102)
- refactor(harness): expand independent trace coverage by [@​yujonglee-berri](https://github.com/yujonglee-berri) in [#​41120](https://github.com/BerriAI/litellm/pull/41120)
- fix(proxy): release max\_parallel\_requests slot when a realtime session ends without LLM callbacks by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41113](https://github.com/BerriAI/litellm/pull/41113)
- fix(router): cool down team deployments on 429 when a sibling serves the same public model by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40991](https://github.com/BerriAI/litellm/pull/40991)
- fix(ui): move tags typed into key metadata JSON into the Tags field by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41023](https://github.com/BerriAI/litellm/pull/41023)
- fix(ui): let team admins grant a team all proxy models by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​40196](https://github.com/BerriAI/litellm/pull/40196)
- chore(lint): graduate 12 rules from the strict-gate ratchet by [@​HUAHAODIA](https://github.com/HUAHAODIA) in [#​41048](https://github.com/BerriAI/litellm/pull/41048)
- test: add dedicated CircleCI integration contract foundation by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41066](https://github.com/BerriAI/litellm/pull/41066)
- fix(utils): keep litellm params out of provider request bodies by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41018](https://github.com/BerriAI/litellm/pull/41018)
- fix(openai): keep extra\_headers out of the chat request body on the httpx handler path by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41141](https://github.com/BerriAI/litellm/pull/41141)
- test: cover persisted updates and warmed authorization policies by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41070](https://github.com/BerriAI/litellm/pull/41070)
- fix(proxy): resolve x-litellm-call-id from response metadata when routes omit call\_id by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41056](https://github.com/BerriAI/litellm/pull/41056)
- chore(prices): sync Vertex AI prices: 14 models by [@​berriai-litellm-provider-info-sync](https://github.com/berriai-litellm-provider-info-sync)\[bot] in [#​40955](https://github.com/BerriAI/litellm/pull/40955)
- ci(codeql): exclude noisy Python quality queries by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41142](https://github.com/BerriAI/litellm/pull/41142)
- feat(proxy): predict prompt-cache costs across deployments by [@​tin-berri](https://github.com/tin-berri) in [#​40877](https://github.com/BerriAI/litellm/pull/40877)
- fix(prompt\_security): keep polling file sanitization through non-terminal statuses by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41131](https://github.com/BerriAI/litellm/pull/41131)
- fix(health): skip background health check DB writes when the latest-row read fails by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41145](https://github.com/BerriAI/litellm/pull/41145)
- feat(model\_armor): logging\_only mode scans completed streams after delivery by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40702](https://github.com/BerriAI/litellm/pull/40702)
- fix(bedrock guardrails): derive contextual grounding source and query from plain messages by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41132](https://github.com/BerriAI/litellm/pull/41132)
- fix(cli): drop enum.StrEnum so the CLI imports on Python 3.10 by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41046](https://github.com/BerriAI/litellm/pull/41046)
- fix(responses): route mid-stream error events through exception\_type so content\_policy\_fallbacks fire by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40988](https://github.com/BerriAI/litellm/pull/40988)
- fix(cost): bill gemini-embedding-2 per token and stop double charging audio by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41157](https://github.com/BerriAI/litellm/pull/41157)
- test: bind management E2E callers and isolate JWT actors by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​40892](https://github.com/BerriAI/litellm/pull/40892)
- fix(headroom): protect cache\_control-marked rows anywhere in history by [@​rad-p44](https://github.com/rad-p44) in [#​40315](https://github.com/BerriAI/litellm/pull/40315)
- test: add strict stateless provider replay identity by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41149](https://github.com/BerriAI/litellm/pull/41149)
- fix(ci): test checked-out model pricing in unit jobs by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41181](https://github.com/BerriAI/litellm/pull/41181)
- fix(guardrails): write per-message guardrail rewrites back onto Responses input items by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40939](https://github.com/BerriAI/litellm/pull/40939)
- fix(proxy): log the provider usage on deferred /v1/messages calls and price cache writes without a creation rate by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41172](https://github.com/BerriAI/litellm/pull/41172)
- fix(guardrails): record not\_run evaluation when scoping leaves nothing to scan by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​39050](https://github.com/BerriAI/litellm/pull/39050)
- fix(responses): preserve provider affinity by [@​AaronHowell](https://github.com/AaronHowell) in [#​40228](https://github.com/BerriAI/litellm/pull/40228)
- fix(sdk): keep body and proxy headers on BadRequestError mapped from a litellm\_proxy 400 by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40994](https://github.com/BerriAI/litellm/pull/40994)
- fix(responses): hoist Codex additional\_tools input items into the chat bridge tools by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40989](https://github.com/BerriAI/litellm/pull/40989)
- fix(router): honor team and key provider weights by [@​tin-berri](https://github.com/tin-berri) in [#​41072](https://github.com/BerriAI/litellm/pull/41072)
- test(e2e): verify streamed answers and tool continuation by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41194](https://github.com/BerriAI/litellm/pull/41194)
- fix(cli): label router costs and simplify the routed-model header by [@​tin-berri](https://github.com/tin-berri) in [#​41186](https://github.com/BerriAI/litellm/pull/41186)
- test(spend): reconcile concurrent requests and daily activity by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41188](https://github.com/BerriAI/litellm/pull/41188)
- fix(guardrails): scan the Anthropic top-level system prompt and tool\_use arguments by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40984](https://github.com/BerriAI/litellm/pull/40984)
- fix(router): count num\_retries\_per\_request across fallback hops by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41191](https://github.com/BerriAI/litellm/pull/41191)
- fix(bedrock): grant rerank, retrieve, agent, and agentcore actions in the web identity session policy by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41168](https://github.com/BerriAI/litellm/pull/41168)
- fix(vertex-live): bill Gemini Live sessions end to end (internal copy of [#​37075](https://github.com/BerriAI/litellm/issues/37075)) by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40915](https://github.com/BerriAI/litellm/pull/40915)
- fix(health): resolve litellm\_credential\_name in realtime health checks by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41173](https://github.com/BerriAI/litellm/pull/41173)
- feat(proxy): unified custom\_key\_policy hook for key generate, update and regenerate by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40921](https://github.com/BerriAI/litellm/pull/40921)
- fix(proxy): enforce custom\_key\_update policy on /key/regenerate by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40695](https://github.com/BerriAI/litellm/pull/40695)
- fix(router): preserve session model choice within each complexity tier by [@​tin-berri](https://github.com/tin-berri) in [#​41174](https://github.com/BerriAI/litellm/pull/41174)
- test(pricing): let synced GovCloud Bedrock rows cite the AWS price list by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41263](https://github.com/BerriAI/litellm/pull/41263)
- docs(github): ask for interactive coding-tool proof in the PR template by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41257](https://github.com/BerriAI/litellm/pull/41257)
- feat(proxy): add POST /management/v1/users/bulk for batched user and team membership creation by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41028](https://github.com/BerriAI/litellm/pull/41028)
- fix(credentials): answer 409 on a credential name collision, make Terraform adoption opt-in by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​40917](https://github.com/BerriAI/litellm/pull/40917)
- feat(proxy): add POST /management/v1/users/bulk\_delete and POST /management/v1/teams/{team\_id}/members/bulk\_delete by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41039](https://github.com/BerriAI/litellm/pull/41039)
- fix(proxy): list directly assigned team models in model access errors by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41256](https://github.com/BerriAI/litellm/pull/41256)
- feat(auto-router): allow opted-in team members to manage their routers by [@​tin-berri](https://github.com/tin-berri) in [#​41175](https://github.com/BerriAI/litellm/pull/41175)
- build(rust-bridge): add typed \_native stub and validate it with mypy.stubtest by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41180](https://github.com/BerriAI/litellm/pull/41180)
- feat(guardrails): add new upstream presidio pii entities including german set by [@​MvdB](https://github.com/MvdB) in [#​36775](https://github.com/BerriAI/litellm/pull/36775)
- fix(responses): filter bridged kwargs like the native Responses path by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41144](https://github.com/BerriAI/litellm/pull/41144)
- test(e2e): cover the reliability retry, cooldown, fallback, and routing-strategy cells by [@​mateo-berri](https://github.com/mateo-berri) in [#​39857](https://github.com/BerriAI/litellm/pull/39857)
- fix(anthropic): add the per-turn-control beta when a message carries output\_config by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41189](https://github.com/BerriAI/litellm/pull/41189)
- fix(router): bind per-request routing\_strategy override selectors to the request's callbacks by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41178](https://github.com/BerriAI/litellm/pull/41178)
- feat(proxy): bind JWT claims to registered agents via agent\_id\_jwt\_field by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40904](https://github.com/BerriAI/litellm/pull/40904)
- fix(proxy): enforce organization budgets when max\_budget is 0 by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​41271](https://github.com/BerriAI/litellm/pull/41271)
- fix(alerting): send llm\_exceptions Slack alert for 5xx HTTPException and ProxyException by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41125](https://github.com/BerriAI/litellm/pull/41125)
- fix(headroom): protect the cached prefix through the last cache\_control breakpoint by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41161](https://github.com/BerriAI/litellm/pull/41161)
- fix(utils): cache custom HuggingFace tokenizers across /utils/token\_counter requests by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41216](https://github.com/BerriAI/litellm/pull/41216)
- fix(router): keep weighted routing when a deployment id equals a model\_name by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41156](https://github.com/BerriAI/litellm/pull/41156)
- feat(router): add capability classifier as Fuse foundation by [@​tin-berri](https://github.com/tin-berri) in [#​41270](https://github.com/BerriAI/litellm/pull/41270)
- fix(proxy): keep access-group raw SQL writes on the writer while writer\_unavailable is stale by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41283](https://github.com/BerriAI/litellm/pull/41283)
- fix(prometheus): count 401 auth failures in litellm\_proxy\_failed\_requests\_metric by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41170](https://github.com/BerriAI/litellm/pull/41170)
- test: drop tests that pin vendor facts and add the CLAUDE.md rule by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41269](https://github.com/BerriAI/litellm/pull/41269)
- fix(proxy): run the remaining inline token counts off the event loop by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40262](https://github.com/BerriAI/litellm/pull/40262)
- fix(proxy): log blocked streaming guardrail responses as failures, not success by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40191](https://github.com/BerriAI/litellm/pull/40191)
- feat(proxy): add tpd\_limit (tokens per day) for batch submissions by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40997](https://github.com/BerriAI/litellm/pull/40997)
- fix(proxy): reconcile budget reservation before enqueuing spend to the DB by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40310](https://github.com/BerriAI/litellm/pull/40310)
- fix(xai): stop sending web\_search\_options to xAI's retired Live Search path by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​38278](https://github.com/BerriAI/litellm/pull/38278)
- feat(terraform): add tpm\_limit, rpm\_limit, budget\_duration, allowed\_models to litellm\_team\_member\_add by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​38682](https://github.com/BerriAI/litellm/pull/38682)
- fix(rerank): bill Vertex search\_units from input records and give every rerank response a unique id by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35180](https://github.com/BerriAI/litellm/pull/35180)
- fix(router): stop counting caller-set timeout 408s toward deployment cooldown by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41230](https://github.com/BerriAI/litellm/pull/41230)
- feat(router): add Fuse V2 classifier after capability forecasting by [@​tin-berri](https://github.com/tin-berri) in [#​41272](https://github.com/BerriAI/litellm/pull/41272)
- fix(proxy): keep client User-Agent on auth failure spend logs by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41291](https://github.com/BerriAI/litellm/pull/41291)
- fix(proxy): reset budgets by decrementing pre-reset spend instead of zeroing rows by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41279](https://github.com/BerriAI/litellm/pull/41279)
- fix(xai): honor nested web\_search filters on the xAI Responses API by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​38268](https://github.com/BerriAI/litellm/pull/38268)
- fix(router): stop registering a caller-supplied credential as a router deployment by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​41289](https://github.com/BerriAI/litellm/pull/41289)
- fix(router): accept custom\_provider\_map providers before the first completion call by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41300](https://github.com/BerriAI/litellm/pull/41300)
- fix(proxy): return 400 instead of 500 for lone surrogate escapes in request body by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41297](https://github.com/BerriAI/litellm/pull/41297)
- fix(langsmith): keep events appended during an in-flight flush instead of clearing them by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41288](https://github.com/BerriAI/litellm/pull/41288)
- fix(logging): track spend for streams a deployment hook converted to non-streaming by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41171](https://github.com/BerriAI/litellm/pull/41171)
- fix(bedrock): sanitize client tool\_call ids to Bedrock toolUseId constraints by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40872](https://github.com/BerriAI/litellm/pull/40872)
- fix(passthrough): attribute Vertex passthrough successes to the resolved router deployment by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41307](https://github.com/BerriAI/litellm/pull/41307)
- feat(ui): persist Models table search, filters, sort and page in the URL by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41296](https://github.com/BerriAI/litellm/pull/41296)
- fix(jwt-auth): scope JWT key mappings by issuer to prevent cross-issuer collisions by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​41281](https://github.com/BerriAI/litellm/pull/41281)
- feat(openai): add openai\_system\_messages\_first to put system messages first for prompt caching by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41304](https://github.com/BerriAI/litellm/pull/41304)
- feat(ui): add custom request headers to the API Playground by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41309](https://github.com/BerriAI/litellm/pull/41309)
- feat(cli): sync Codex /model picker from proxy /v1/models in lite codex by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40476](https://github.com/BerriAI/litellm/pull/40476)
- chore: bump litellm-enterprise 0.1.67 -> 0.1.68, litellm-proxy-extras 0.4.97 -> 0.4.98, litellm 1.102.0 -> 1.103.0 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41321](https://github.com/BerriAI/litellm/pull/41321)
- feat: add aihubmix provider pricing entries by [@​IToSSc](https://github.com/IToSSc) in [#​41179](https://github.com/BerriAI/litellm/pull/41179)
- feat(auto-router): add per-model Fast mode toggle by [@​tin-berri](https://github.com/tin-berri) in [#​41282](https://github.com/BerriAI/litellm/pull/41282)
- fix(proxy): include litellm\_call\_id in LLM API exception logs by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41205](https://github.com/BerriAI/litellm/pull/41205)
- fix(proxy): keep yaml pass-through endpoints visible to auth after db overlay by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41303](https://github.com/BerriAI/litellm/pull/41303)
- fix(proxy): resolve router\_settings.model\_group\_alias before key/team model auth by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41308](https://github.com/BerriAI/litellm/pull/41308)
- fix(ui): block usage export and flag the range when a spend page fails by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​41294](https://github.com/BerriAI/litellm/pull/41294)
- fix(vertex\_ai): bill Gemini Omni Interactions usage and Veo sampleCount on passthrough by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41322](https://github.com/BerriAI/litellm/pull/41322)
- fix(proxy): honor LITELLM\_LOG for uvicorn and proxy extras loggers by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41306](https://github.com/BerriAI/litellm/pull/41306)
- fix(cost): price native Responses WebSocket turns at their returned service\_tier by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41318](https://github.com/BerriAI/litellm/pull/41318)
- fix(proxy): key model rpm/tpm override takes precedence over team model limit by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41302](https://github.com/BerriAI/litellm/pull/41302)
- fix(proxy): track per-member organization spend by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41255](https://github.com/BerriAI/litellm/pull/41255)
- feat(proxy): add /nvidia\_nim passthrough route for NIM object detection and OCR /v1/infer by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41316](https://github.com/BerriAI/litellm/pull/41316)
- feat(model\_info): add provider-neutral Gemini 2.5+ chat baseline fallback generalization by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41320](https://github.com/BerriAI/litellm/pull/41320)
- fix(spend): sum multi-round session duration in logs UI by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35388](https://github.com/BerriAI/litellm/pull/35388)
- feat(router): limit unlicensed Capability and Fuse v2 routers to one each by [@​tin-berri](https://github.com/tin-berri) in [#​41326](https://github.com/BerriAI/litellm/pull/41326)
- fix(guardrails): resolve caller identity from metadata buckets in custom code guardrail by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41126](https://github.com/BerriAI/litellm/pull/41126)
- fix(e2e): onboard dashboard users through invitations by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41319](https://github.com/BerriAI/litellm/pull/41319)
- feat(guardrails): add Microsoft Agent 365 MCP tool-call guardrail by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​38241](https://github.com/BerriAI/litellm/pull/38241)
- test: drop remaining tests that pin cost-map vendor facts by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41298](https://github.com/BerriAI/litellm/pull/41298)
- feat(ui): show average response time per model in usage model activity by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41313](https://github.com/BerriAI/litellm/pull/41313)
- fix(proxy): preserve Anthropic pricing modifiers in router savings by [@​tin-berri](https://github.com/tin-berri) in [#​41341](https://github.com/BerriAI/litellm/pull/41341)
- feat(guardrails): support pre\_call and during\_call modes for llm\_as\_a\_judge by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41128](https://github.com/BerriAI/litellm/pull/41128)
- fix(gemini): propagate the provider's modelVersion to the response model by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41338](https://github.com/BerriAI/litellm/pull/41338)
- fix(fireworks-ai): bill cache-write, reasoning and audio tokens via the shared cost calculator by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41339](https://github.com/BerriAI/litellm/pull/41339)
- feat(guardrails): singulr v2 API contract with logging\_only, pre\_mcp\_call and post\_mcp\_call by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​41329](https://github.com/BerriAI/litellm/pull/41329)
- ci(image-scan): ignore zlib CVE-2026-85091 until Wolfi ships the fix by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41353](https://github.com/BerriAI/litellm/pull/41353)
- feat(e2e): reuse exact provider responses for 24 hours by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41346](https://github.com/BerriAI/litellm/pull/41346)
- fix(xai): keep 'instructions' on the xAI Responses API so system messages survive web search by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​38254](https://github.com/BerriAI/litellm/pull/38254)
- feat(ui): configure capability and Fuse v2 classifiers by [@​tin-berri](https://github.com/tin-berri) in [#​41315](https://github.com/BerriAI/litellm/pull/41315)
- fix(anthropic): tolerate message\_delta events without usage when streaming by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41336](https://github.com/BerriAI/litellm/pull/41336)
- test(router): ignore deployment-selection logs in the fallback log assertion by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41358](https://github.com/BerriAI/litellm/pull/41358)
- test(proxy): assert budget resets decrement the cleared spend by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41359](https://github.com/BerriAI/litellm/pull/41359)
- fix(e2e): expect models filters to persist after reload by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41348](https://github.com/BerriAI/litellm/pull/41348)
- fix(e2e): record cookie-setting provider responses and keep prompt-caching tests live by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41366](https://github.com/BerriAI/litellm/pull/41366)
- fix(responses): recount tokens when a streamed response completes without usage by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41337](https://github.com/BerriAI/litellm/pull/41337)
- fix(ui): simplify Capability and Fuse advanced routing options by [@​tin-berri](https://github.com/tin-berri) in [#​41371](https://github.com/BerriAI/litellm/pull/41371)
- fix(mcp): authorize JWT OAuth credential persistence by [@​joshua-berri](https://github.com/joshua-berri) in [#​41314](https://github.com/BerriAI/litellm/pull/41314)
- feat(router): stream shadow traffic and fan out silent\_model to multiple targets by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41368](https://github.com/BerriAI/litellm/pull/41368)
- perf(content\_filter): scan a bounded window per streamed chunk by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41407](https://github.com/BerriAI/litellm/pull/41407)
- fix(proxy): hide model allowlist from client-facing model access denied errors by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41310](https://github.com/BerriAI/litellm/pull/41310)
- feat(http): opt-in outbound HTTP/2 for httpx clients by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41268](https://github.com/BerriAI/litellm/pull/41268)
- refactor(rust): remove gateway, config, router, realtime, and Rust trace-parity instrumentation by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41432](https://github.com/BerriAI/litellm/pull/41432)
- fix(guardrails): don't add post\_call output scan for MCP-only Presidio modes by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40571](https://github.com/BerriAI/litellm/pull/40571)
- chore(prices): sync Azure, Azure AI, Gemini, OpenAI, Bedrock, Together AI, Fireworks and Vertex prices: 278 models, 59 new, 30 deprecated by [@​berriai-litellm-provider-info-sync](https://github.com/berriai-litellm-provider-info-sync)\[bot] in [#​41154](https://github.com/BerriAI/litellm/pull/41154)
- fix(rag): forward retrieval\_filter from retrieval\_config to vector store search by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34427](https://github.com/BerriAI/litellm/pull/34427)
- refactor(rust): extract auth and cache crates by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41464](https://github.com/BerriAI/litellm/pull/41464)
- fix(proxy): default litellm\_trace\_id to the OTel server span trace id by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41386](https://github.com/BerriAI/litellm/pull/41386)
- chore(codeowners): add ryan and kerry as owners of the cost map by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41333](https://github.com/BerriAI/litellm/pull/41333)
- fix(responses): guard empty-choices chunks in the Responses API streaming bridge by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34455](https://github.com/BerriAI/litellm/pull/34455)
- chore(prices): sync Google Gemini prices: 22 models by [@​berriai-litellm-provider-info-sync](https://github.com/berriai-litellm-provider-info-sync)\[bot] in [#​41457](https://github.com/BerriAI/litellm/pull/41457)
- fix(bedrock): forward userContext in Knowledge Base Retrieve requests by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41475](https://github.com/BerriAI/litellm/pull/41475)
- ci(rust): split rust jobs, use nextest and Swatinem/rust-cache by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41480](https://github.com/BerriAI/litellm/pull/41480)
- fix(fireworks\_ai): flatten dict-form reasoning\_effort to its effort string by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41335](https://github.com/BerriAI/litellm/pull/41335)
- fix(proxy): never forward the LiteLLM virtual key to Anthropic on the /anthropic passthrough by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41340](https://github.com/BerriAI/litellm/pull/41340)
- fix(proxy): rename AWS Secrets Manager secret when key alias changes by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41468](https://github.com/BerriAI/litellm/pull/41468)
- feat(otel): promote nested request metadata keys to litellm.metadata.\* span attributes by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41462](https://github.com/BerriAI/litellm/pull/41462)
- fix(proxy): sync AWS Secrets Manager on body-less key regenerate and key alias changes by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41458](https://github.com/BerriAI/litellm/pull/41458)
- fix(http\_handler): keep a handler alive while a response it issued is still reading by [@​max-sixty](https://github.com/max-sixty) in [#​34829](https://github.com/BerriAI/litellm/pull/34829)
- fix(bedrock): make prompt caching work on the Nova InvokeModel route by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41343](https://github.com/BerriAI/litellm/pull/41343)
- ci(migrations): flag defaulted ADD COLUMN on request-log tables by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41460](https://github.com/BerriAI/litellm/pull/41460)
- feat(prometheus): add customer (end\_user) budget gauges by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41472](https://github.com/BerriAI/litellm/pull/41472)
- fix(otel): drop None metric and event attributes before OTLP export by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36815](https://github.com/BerriAI/litellm/pull/36815)
- fix(anthropic): carry the served model from message\_start onto stream chunks by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41446](https://github.com/BerriAI/litellm/pull/41446)
- fix(models): rolling registry audit: Gemini latest aliases, Nova cache pricing, OpenRouter/Together sync, Mistral GLM 5.3, Azure snapshots, Grok caching by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41112](https://github.com/BerriAI/litellm/pull/41112)
- fix(router): count TPM/RPM usage before building rate-limit headers by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41474](https://github.com/BerriAI/litellm/pull/41474)
- feat(guardrails): release buffered stream chunks after each passing scan by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41425](https://github.com/BerriAI/litellm/pull/41425)
- fix!: re-check budget on router fallback targets by [@​runjivu](https://github.com/runjivu) in [#​41379](https://github.com/BerriAI/litellm/pull/41379)
- refactor(ocr): move file preparation from the python bridge into litellm-core by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41489](https://github.com/BerriAI/litellm/pull/41489)
- feat(s3): add s3\_log\_prompts\_only option to log prompts without responses by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41327](https://github.com/BerriAI/litellm/pull/41327)
- feat(team): team-level model\_max\_budget with key-level overrides by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41330](https://github.com/BerriAI/litellm/pull/41330)
- feat(keys): filter /key/list by active, expired, revoked or deleted status and serve deleted keys from /key/info by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41311](https://github.com/BerriAI/litellm/pull/41311)
- feat(proxy): expose lifetime total\_spend on virtual keys by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41403](https://github.com/BerriAI/litellm/pull/41403)
- fix(proxy): release completed max-parallel slots promptly by [@​elifozdamar](https://github.com/elifozdamar) in [#​40843](https://github.com/BerriAI/litellm/pull/40843)
- feat(ui): accept ssh clone urls when registering a skill by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35418](https://github.com/BerriAI/litellm/pull/35418)
- fix(prices): dedupe Nova cache\_read\_input\_token\_cost keys left by a text merge by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41496](https://github.com/BerriAI/litellm/pull/41496)
- fix(otel): propagate W3C trace context on HTTP and WebSocket passthrough by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40669](https://github.com/BerriAI/litellm/pull/40669)
- fix(proxy): remove duplicate user budget hook that 429'd zero-cost models by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41345](https://github.com/BerriAI/litellm/pull/41345)
- test(logging): pick this test's own records out of the shared log batch by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41487](https://github.com/BerriAI/litellm/pull/41487)
- test(together\_ai): move request-shape checks to the mapped file, drop the live ones by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41360](https://github.com/BerriAI/litellm/pull/41360)
- feat(ui): shared URL-state layer for tables and tabs by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​41331](https://github.com/BerriAI/litellm/pull/41331)
- feat(e2e): make the provider cache reusable across builds and mount Bedrock behind it by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41402](https://github.com/BerriAI/litellm/pull/41402)
- fix(mcp): fail closed on missing upstream credentials by [@​joshua-berri](https://github.com/joshua-berri) in [#​41364](https://github.com/BerriAI/litellm/pull/41364)
- feat(rust): scaffold Redis cache crate by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41501](https://github.com/BerriAI/litellm/pull/41501)
- fix(dashscope): forward reasoning\_effort to the provider by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​37506](https://github.com/BerriAI/litellm/pull/37506)
- fix(proxy): carry litellm\_call\_id through endpoint specific error logs and failure responses by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41356](https://github.com/BerriAI/litellm/pull/41356)
- fix(proxy): retry rate-limit fallbacks from a pristine request snapshot by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40596](https://github.com/BerriAI/litellm/pull/40596)
- fix(gemini): map minimal thinking to low for Gemini 3.7 and 3.8 Flash by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41201](https://github.com/BerriAI/litellm/pull/41201)
- fix(proxy): stop forwarding LiteLLM credential headers on Bedrock agent-runtime passthrough by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41504](https://github.com/BerriAI/litellm/pull/41504)
- fix(streaming): estimate interrupted Anthropic stream usage from reasoning\_content by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41503](https://github.com/BerriAI/litellm/pull/41503)
- fix(azure\_ai): route Responses API to native /openai/v1/responses for Foundry Models by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33856](https://github.com/BerriAI/litellm/pull/33856)
- fix(proxy): show all model groups to proxy admins in /model\_group/info by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41094](https://github.com/BerriAI/litellm/pull/41094)
- feat(proxy): let proxy admins choose which team fields team admins may edit by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​39996](https://github.com/BerriAI/litellm/pull/39996)
- fix(bedrock): neutralize orphaned tool blocks instead of raising or injecting a dummy tool (internal copy of [#​31400](https://github.com/BerriAI/litellm/issues/31400)) by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41513](https://github.com/BerriAI/litellm/pull/41513)
- feat(ui): persist organizations and projects list, detail tab and key table state in the URL by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​41445](https://github.com/BerriAI/litellm/pull/41445)
- fix(bedrock\_mantle): accept and forward verbosity on gpt-5.x chat completions by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41509](https://github.com/BerriAI/litellm/pull/41509)
- test: cover database transactions and persisted accounting by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41073](https://github.com/BerriAI/litellm/pull/41073)
- test: provider wire contracts, streaming and recovery by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41075](https://github.com/BerriAI/litellm/pull/41075)
- fix(mcp): count admin static headers as api\_key credential slots by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41514](https://github.com/BerriAI/litellm/pull/41514)
- ci: auto-merge provider-info-sync PRs when CI, Greptile and Bugbot are clean by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41494](https://github.com/BerriAI/litellm/pull/41494)
- feat(rust): add standalone framing crate by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41500](https://github.com/BerriAI/litellm/pull/41500)
- fix(proxy): enforce tag budgets for tags added by guardrails by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40842](https://github.com/BerriAI/litellm/pull/40842)
- fix(utils): run post-call deployment hook on converted chat streams by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41495](https://github.com/BerriAI/litellm/pull/41495)
- fix(e2e): bind provider-cache recordings to the deployment's test, not the serving process by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41520](https://github.com/BerriAI/litellm/pull/41520)
- test: add extension and browser integration contracts by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41078](https://github.com/BerriAI/litellm/pull/41078)
- fix(logging): scan each log record once and collapse base64 payloads before the secret regex by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40934](https://github.com/BerriAI/litellm/pull/40934)
- fix(spend\_tracking): attribute router-rejected requests to the model group provider by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41507](https://github.com/BerriAI/litellm/pull/41507)
- feat(router): discover token limits for hosted OpenAI-compatible models by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41508](https://github.com/BerriAI/litellm/pull/41508)
- feat(proxy): let team admins edit rpm\_limit and max\_budget when enabled by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​41525](https://github.com/BerriAI/litellm/pull/41525)
- fix(otel): fit per-index OpenInference messages to the span's remaining attribute budget by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41498](https://github.com/BerriAI/litellm/pull/41498)
- test(aws): verify rotated secret value by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41524](https://github.com/BerriAI/litellm/pull/41524)
- test(e2e): read a deleted key back as deleted, not as a 404 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41551](https://github.com/BerriAI/litellm/pull/41551)
- fix(otel v2): map the caller's Langfuse user, session and tags onto the root and generation spans by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41140](https://github.com/BerriAI/litellm/pull/41140)
- test: fix seven tests left stale by [#​41311](https://github.com/BerriAI/litellm/issues/41311), [#​41337](https://github.com/BerriAI/litellm/issues/41337), [#​39996](https://github.com/BerriAI/litellm/issues/39996), [#​41310](https://github.com/BerriAI/litellm/issues/41310), [#​41289](https://github.com/BerriAI/litellm/issues/41289) and [#​41315](https://github.com/BerriAI/litellm/issues/41315) by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41527](https://github.com/BerriAI/litellm/pull/41527)
- test(budgets): cover management null handling by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41563](https://github.com/BerriAI/litellm/pull/41563)
- test(e2e): drop the auto-router select "opens below" spec by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41568](https://github.com/BerriAI/litellm/pull/41568)
- fix(guardrails): stream Prompt Security post\_call redactions in incremental\_diff mode by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41558](https://github.com/BerriAI/litellm/pull/41558)
- fix(guardrails): give post-call scans the scoped request conversation and tools by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41220](https://github.com/BerriAI/litellm/pull/41220)
- feat(openrouter): add stealth/union-alpha to the model cost map by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41576](https://github.com/BerriAI/litellm/pull/41576)
- test(management): cover project authorization lifecycle by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41573](https://github.com/BerriAI/litellm/pull/41573)
- feat(rust): map Anthropic Messages transformations by [@​yujonglee-berri](https://github.com/yujonglee-berri) in [#​41531](https://github.com/BerriAI/litellm/pull/41531)
- fix(e2e): clear the three standing errors in the scheduled Buildkite suite by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41616](https://github.com/BerriAI/litellm/pull/41616)
- refactor(rust\_bridge): declarative route catalog and shared runtime selection by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41479](https://github.com/BerriAI/litellm/pull/41479)
- fix(mock\_completion): keep the resolved provider so router custom pricing resolves for azure\_ai deployments by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41623](https://github.com/BerriAI/litellm/pull/41623)
- fix(mcp): restrict health discovery to virtual key grants by [@​joshua-berri](https://github.com/joshua-berri) in [#​41609](https://github.com/BerriAI/litellm/pull/41609)
- fix(mcp): preserve request-selected guardrails during tool execution by [@​joshua-berri](https://github.com/joshua-berri) in [#​41619](https://github.com/BerriAI/litellm/pull/41619)
- refactor(ocr): mirror Python provider layout and preserve tests by [@​yujonglee-berri](https://github.com/yujonglee-berri) in [#​41550](https://github.com/BerriAI/litellm/pull/41550)
- test(fireworks\_ai): stop pinning vision support on minimax-m3 by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41627](https://github.com/BerriAI/litellm/pull/41627)
- perf(spend\_tracking): index LiteLLM\_SpendLogs by (api\_key, startTime) by [@​etiennechabert](https://github.com/etiennechabert) in [#​37983](https://github.com/BerriAI/litellm/pull/37983)
- fix(proxy): reject non-string model with 400 and log its spend as unknown-model by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41633](https://github.com/BerriAI/litellm/pull/41633)
- test(together\_ai): stop pinning successor deprecation status by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41635](https://github.com/BerriAI/litellm/pull/41635)
- chore(prices): sync Together AI prices: 6 models, 6 deprecated \[sync failed: Google Gemini] by [@​berriai-litellm-provider-info-sync](https://github.com/berriai-litellm-provider-info-sync)\[bot] in [#​41570](https://github.com/BerriAI/litellm/pull/41570)
- fix(budgets): page end-user cache invalidation after a budget reset by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​41488](https://github.com/BerriAI/litellm/pull/41488)
- chore: bump litellm-proxy-extras 0.4.98 -> 0.4.99 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41659](https://github.com/BerriAI/litellm/pull/41659)
- fix(tests): resolve the integration support package without run.py's PYTHONPATH by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41373](https://github.com/BerriAI/litellm/pull/41373)
- fix(ui): keep untimed guardrail entries on the request lifecycle by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41374](https://github.com/BerriAI/litellm/pull/41374)
- fix(anthropic-bridge): convert mid-conversation system turns to user turns on /v1/messages to chat completions by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41493](https://github.com/BerriAI/litellm/pull/41493)
- fix(bedrock): support aws-sdk-bedrock-runtime 0.10/0.11 in Bedrock Realtime by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41542](https://github.com/BerriAI/litellm/pull/41542)
- feat(cli): deprecate the litellm-proxy entrypoint in favour of lite by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41673](https://github.com/BerriAI/litellm/pull/41673)
- fix(scim): align pagination `count` validation with RFC 7644 by [@​zachbernstein-sdx](https://github.com/zachbernstein-sdx) in [#​41444](https://github.com/BerriAI/litellm/pull/41444)
- fix(bedrock): never emit Converse cachePoint for OpenAI-family models by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41419](https://github.com/BerriAI/litellm/pull/41419)
- fix(images): stop forwarding the raw image\[] and mask\[] form keys by [@​mateo-berri](https://github.com/mateo-berri) in [#​39512](https://github.com/BerriAI/litellm/pull/39512)
- feat(management\_v1): bulk update team member budgets by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​41632](https://github.com/BerriAI/litellm/pull/41632)
- refactor(rust): extract provider translations by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41690](https://github.com/BerriAI/litellm/pull/41690)
- feat(cli): rename lite autoroute up/down to start/stop, keeping the old names as deprecated aliases by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41672](https://github.com/BerriAI/litellm/pull/41672)
- fix(responses): keep the addressed response id off bridged provider requests by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41689](https://github.com/BerriAI/litellm/pull/41689)
- fix(license): let a wildcard allowed\_features license grant the auto\_router feature by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41684](https://github.com/BerriAI/litellm/pull/41684)
- fix(ui): list every provider in the cache leakage by-model table by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40875](https://github.com/BerriAI/litellm/pull/40875)
- fix(team): keep a forked member budget's reset window and audit bulk member budget writes by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​41686](https://github.com/BerriAI/litellm/pull/41686)
- feat(proxy): add TypeSafe AI Jev evaluate passthrough with registry-priced spend tracking by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41607](https://github.com/BerriAI/litellm/pull/41607)
- test(e2e): cover bedrock batch file upload and create in the us-gov-west-1 partition by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41536](https://github.com/BerriAI/litellm/pull/41536)
- feat(grafana): add all-metrics dashboard and fix stale dashboard\_v2 gauges by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41578](https://github.com/BerriAI/litellm/pull/41578)
- fix(cost): price Azure PTU spillover requests at standard token rates by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41569](https://github.com/BerriAI/litellm/pull/41569)
- build(deps): bump soupsieve to 2.9.2 to clear the osv-scan advisories by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41703](https://github.com/BerriAI/litellm/pull/41703)
- fix(fireworks\_ai): restore supports\_vision on minimax-m3 in the cost map by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41699](https://github.com/BerriAI/litellm/pull/41699)
- feat(policy\_engine): explicit priority for policy attachment execution order by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41571](https://github.com/BerriAI/litellm/pull/41571)
- feat(router): add TypeSafe Jev as a complexity router classifier by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41615](https://github.com…
TLDR
Problem this solves:
mode: logging_only,pre_mcp_callandpost_mcp_callare not available for Singulrnullverdict into an allowHow it solves it:
should_blockis a bool with default false, sonullis an invalid verdictUser Flow
Before: an admin runs Singulr in
mode: logging_onlyand sees wrong or missing verdicts in the logs, and MCP tool calls are scanned as if they were chat promptsmode: logging_only,default_on: trueand restarts the proxy{"name": "echo", "arguments": {"text": "...SSN..."}, "server_id": "..."}; Singulr receives a chatrequestwith the tool argument as a user message instead of an MCP tool call{"should_block": null}underblock_on_error: true; the request goes throughAfter: the same admin sees a request and a response verdict on every call, outages are recorded, and MCP tool calls reach Singulr as MCP tool calls
mode: logging_only,default_on: trueand restarts the proxyguardrail_intervened,Blocking due to Prompt injection detected) withguardrail_mode: logging_only; on a clean prompt it shows one request entry and one response entry, and Singulr evaluates the response as{role, content, tool_calls}guardrail_failed_to_respondentrymcp_requestwithtool_nameandtool_arguments, and the tool result is scanned asmcp_response(the result with the SSN comes back as 400){"should_block": null}underblock_on_error: true; the request is blocked with 400Singulr API returned an invalid responseRelevant issues
Credit
Adopted from #37464 by @aniket-kardile. Mirrored onto a
litellm_branch so CircleCI and the internal lint workflow run, with the fixes listed under Behavior changes on top of the original commitsLinear ticket
Resolves LIT-6623
Behavior changes
Removed / renamed
SingulrGuardrail.async_logging_hookandlogging_hookoverrides from feat(guardrails): update Singulr guardrail api contract and add logging, pre-MCP, and post-MCP hooks #37464 are gone;CustomGuardrail.async_logging_hooknow servesmode: logging_onlySingulrGuardrailResponse.should_blockisbool = False(wasbool | None = Nonein feat(guardrails): update Singulr guardrail api contract and add logging, pre-MCP, and post-MCP hooks #37464);SingulrGuardrailPayload.responseisAssistantMessage | None,messages/toolsareSequence[Mapping[str, object]],metadataisMapping[str, str]andSingulrMcpGuardrailPayload.tool_argumentsisobject(all wereAny-based); a non-dict RESTargumentsvalue is forwarded instead of raising a 500Silent regressions to watch
mode: logging_onlyrecords twoguardrail_informationentries per successful call (request and response), each withguardrail_modeandguardrail_response; feat(guardrails): update Singulr guardrail api contract and add logging, pre-MCP, and post-MCP hooks #37464 recorded one entry per call. Anything counting entries per request sees 2logging_onlyends the scan; the response leg is not sent (matches enforcement, where a blocked request has no response){"should_block": null}from Singulr is an invalid verdict: 400 underblock_on_error: true, allow with an error log underblock_on_error: false. feat(guardrails): update Singulr guardrail api contract and add logging, pre-MCP, and post-MCP hooks #37464 treated it as allow; the v1 contract onmainfails closed the same way this PR does/mcp-rest/tools/callpre-scans reach Singulr asmcp_request; feat(guardrails): update Singulr guardrail api contract and add logging, pre-MCP, and post-MCP hooks #37464 sent the first of the two pre-scans as a chatrequest/v1/embeddingsundermode: logging_onlyis not scanned, the same as everyapply_guardrailguardrail onmaintoday; feat(guardrails): update Singulr guardrail api contract and add logging, pre-MCP, and post-MCP hooks #37464's own logging hook happened to send the embeddings input plus a vector dump because it forwarded every raw callback result. Gating that scan oninspect_embeddingsis a follow-upRegression tests
tests/test_litellm/proxy/guardrails/guardrail_hooks/test_singulr.py::TestSingulrLoggingHookpins the base-path scan (assistant-message response shape, block recorded asguardrail_intervenedwithguardrail_mode: logging_only, vendor timeout recorded asguardrail_failed_to_respond, MCP tool results asmcp_response, no sync hook)::TestSingulrMcpRequest::test_llm_request_body_keys_cannot_reroute_the_scan_to_mcpand::test_llm_response_with_spoofed_mcp_keys_still_scans_the_tool_callspin thatname,arguments,mcp_tool_name,call_typein a chat body never change the scan::TestSingulrMcpRequest::test_mcp_rest_body_shape_routes_to_mcp_request_payloadand::test_mcp_rest_body_without_arguments_still_routes_to_mcp_requestpin the REST body under the proxy'scall_mcp_toollogger::TestSingulrMcpResponse::test_mcp_response_is_detected_from_each_producerpins post_mcp_call, logging_only and REST producers::TestSingulrAllowAction::test_explicit_null_verdict_fails_closed_by_defaultand::test_explicit_null_verdict_fails_open_when_block_on_error_falsepin the null verdictPre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Screenshots / Proof of Fix
Final audit run on the PR tip 1460856 (every id, screenshot and number below is from that hash). Rig: proxy from each commit on its own port with
--num_workers 4, embedded Postgres, real OpenAIgpt-4.1-mini/text-embedding-3-small, real Anthropic, real Singulr demo tenant behind a transparent recording proxy on 127.0.0.1:15735 (the "Singulr received" lines below are its recordings). Before is #37464's three files onmain(8cce2b1, what the original PR would ship);mainitself has the v1 contract, nologging_onlyoverride and no MCP hooks, so a comparison against it only shows the feature is newDashboard (tip 1460856,
mode: logging_only, real Singulr)Screen recording of the whole flow (login, Guardrails page, Logs filtered by id, Guardrails & Policy section for the PII call and the clean call, Playground PII prompt, Logs after): ui-flow-146085669c.webm
Guardrails page listing
singulr-logPII prompt, request
chatcmpl-EOdy1SNnzQUViEQSBS51Gq7aQ6wwM: status Success (logging_only never blocks), Guardrails & Policy shows 1 guardrail evaluated,singulr-log LOGGING-ONLY FAILEDat 2343 ms for the request scope and no response legClean prompt, request
chatcmpl-EOdy2061fUpYAmidcNVVikdnvHa0Q: 2 guardrails evaluated, twosingulr-log LOGGING-ONLY PASSEDentries, request scope 2440 ms and response scope 664 msPlayground, PII prompt through
gpt-4.1-miniunder logging_only: the answer is delivered and the spend rowchatcmpl-EOdz1NPPUfa97V5OYbpno8BY6GfHwcarriessingulr-log / logging_only / guardrail_intervenedLogs after the Playground call; the
/v1/embeddingsrow carries no guardrail entry, the same asmaintoday (see Caveats)Shared config and payloads
modeswapped per case:Before (8cce2b1, #37464 on main)
logging_only, prompt with PII and injection
curl -s -X POST http://127.0.0.1:15736/v1/chat/completions -H "Authorization: Bearer sk-1234" -H "Content-Type: application/json" -d "{\"model\":\"gpt-4.1-mini\",\"max_tokens\":40,\"messages\":[{\"role\":\"user\",\"content\":\"$PII\"}]}"chatcmpl-EOdcT00lJkooKXWa5wHXbjZchqtmL; spend rowguardrail_information: 1 entryguardrail_status: guardrail_intervenedguardrail_scope: request(blocked, 2650 ms) andguardrail_scope: responsecarrying the raw{choices, created, id, model, object, system_fingerprint, usage}dump, answered in 115 ms with no evaluationlogging_only, clean prompt, streamed
curl -s -N -X POST http://127.0.0.1:15736/v1/chat/completions -H "Authorization: Bearer sk-1234" -H "Content-Type: application/json" -d "{\"model\":\"gpt-4.1-mini\",\"max_tokens\":60,\"stream\":true,\"messages\":[{\"role\":\"user\",\"content\":\"$BENIGN\"}]}"chatcmpl-EOdcVLjY03LFKb1NM9ecAlr6qLvek; spend row: 1 entrysuccessModelResponsedump, answered in 139 mslogging_only, Singulr unreachable (socket that never answers, timeout 3 s)
singulr_api_base: http://127.0.0.1:19097chatcmpl-EOde8JNjHSFcU48TFZqi8szySlDEF; spend rowguardrail_information: empty, the outage leaves no tracelogging_only with a legacy sync
success_callbackconfiguredlitellm_settings.success_callback: [hook_recorder.sync_success]plus the same clean requestchatcmpl-EOdd9qFxVOxzujmhruDarTTkdpBDB; spend row: 2 entries,guardrail_intervenedthensuccess: the sync hook ran_call_apifrom the callback thread against the shared client and recorded a spurious blocklogging_only, /v1/embeddings
curl -s -X POST http://127.0.0.1:15736/v1/embeddings -H "Authorization: Bearer sk-1234" -H "Content-Type: application/json" -d "{\"model\":\"text-embedding-3-small\",\"input\":\"$BENIGN\"}"responsescan of the embedding vector dump, answered in 234 mspre_mcp_call + post_mcp_call, REST tool call with PII in the arguments
curl -s -X POST http://127.0.0.1:15736/mcp-rest/tools/call -H "Authorization: Bearer sk-1234" -H "Content-Type: application/json" -d "{\"name\":\"prrisk_tools-echo\",\"arguments\":{\"text\":\"$PII\"},\"server_id\":\"<id>\"}"Guardrail: singulr-mcp-pre ... Blocking due to Prompt injection detectedguardrail_scope: requestwithmessages: [{"role": "user", "content": "<the SSN text>"}], the tool argument dressed as a chat prompt; nomcp_requestscan for this callpre_call,
block_on_error: true, Singulr answers{"should_block": null}curl -s -X POST http://127.0.0.1:15736/v1/chat/completions ... -d "{\"model\":\"gpt-4.1-mini\",\"guardrails\":[\"snullblock\"],\"messages\":[{\"role\":\"user\",\"content\":\"$BENIGN\"}]}"withsingulr_api_basepointed at a stub answering{"should_block": null}chatcmpl-EOdf9wPt9mEwmaKoMXcJNxieEtMYV; spend row:snullblock: successAfter (1460856)
logging_only, prompt with PII and injection
chatcmpl-EOZYCa4dmGdJyOIlnJRcKDxaRNAUW; spend rowguardrail_information: 1 entryguardrail_status: guardrail_intervened,guardrail_mode: logging_only,guardrail_response: "... Blocking due to Prompt injection detected"guardrail_scope: request(blocked, 1713 ms); the response leg is not sent after a request blocklogging_only, clean prompt, streamed
chatcmpl-EOZYEwlz2ULB2GcSKL6tOwa9l3eN1; spend row: 2 entries,success(request) andsuccess(response){"role": "assistant", "content": "Paris", "tool_calls": []}, evaluated in 1334 ms; a tool-call response (chatcmpl-EOZYHsNblrBZJAjgPkpzyV4R2tYqR) arrives as{"role": "assistant", "content": null, "tool_calls": [{"id": "call_...", "type": "function", "function": {"name": "get_weather", "arguments": "{\"city\":\"Paris\"}"}}]}/v1/messages(msg_011Cf6KUkmkkcGY4KvoynimE) and/v1/responses(resp_FgpXMoR2jBIDsi2G...) also get both legs: request 1663 ms plus response 691 ms, and request 1993 ms plus response 1723 mslogging_only, Singulr unreachable (socket that never answers, timeout 3 s)
chatcmpl-EOdjHASX5CrZ3HMVmyu9fUyTy3x6M; spend rowguardrail_information: 1 entryguardrail_status: guardrail_failed_to_respondlogging_only with a legacy sync
success_callbackconfiguredchatcmpl-EOZYrn3d5HVVQGjID0CpP4cbJssEj; spend row: 2 entries,successandsuccess; the callback thread makes no Singulr calllogging_only, /v1/embeddings
OpenAI Embeddings: Unexpected input list item typeand skips, exactly as onmaintoday. feat(guardrails): update Singulr guardrail api contract and add logging, pre-MCP, and post-MCP hooks #37464's scan here came from its raw callback dump (an evaluated request leg plus an unevaluated vector dump); this PR does not carry that over, see Caveatspre_mcp_call + post_mcp_call, REST tool call with PII in the arguments
Guardrail: singulr-mcp-post ... Blocking due to Prompt injection detectedmcp_request(tool_name: prrisk_tools-echo,tool_arguments: {"text": "<the SSN text>"}),mcp_requestagain from the MCP manager's own pre-check (tool_name: echo, generic double run), thenmcp_responsewithtool_result: ["echo: <the SSN text>"], which the tenant blocks (its policy does not enforcemcp_request)pre_call,
block_on_error: true, Singulr answers{"should_block": null}Guardrail: snullblock, Message: Singulr API returned an invalid response: 1 validation error for SingulrGuardrailResponse should_block Input should be a valid boolean; spend row:snullblock: guardrail_intervenedFull matrix and chaos on 1460856
101 cells: 54 happy/sad/edge (pre_call, post_call, key and team level, precedence, bad model, bad provider key, oversized input,
/guardrails/apply_guardrail), 13 logging_only, 1 legacy sync callback, 2 logging_only fault, 6 MCP, 14 fault matrix, 11 DB-stored guardrail across 4 workers. 101 PASS; the DB leg waits for the proxy's periodic guardrail sync before asserting (listed on every worker at once, enforced on every worker after 88 s on this run)Chaos on 4 workers: Singulr killed during a 24-request mixed burst (24 x 400 during the outage, liveliness and unguarded routes 200, every id one spend row, 200 after restart); Singulr paused 8 s under 10 concurrent calls (10 x 200 in 9.1 to 11.0 s, no deadlock, no duplicate rows); proxy killed mid-burst (20 in flight lost, 0 duplicates, 200 after relaunch); one worker killed under traffic (10 x 200, worker respawned)
Base vs head, 36 cells on identical config and order: 18 identical, 18 differ; 17 are head PASS with base FAIL, and one is the intended two-entries-per-call change; no cell is head FAIL with base PASS. The differing cells are the logging_only response shape and repeat counts (L02 to L07, L03b, L05b, L09 x3), the legacy sync callback (L10), the outage record (L12), the MCP scopes (M02 to M04), the null verdict (F04), and
/v1/embeddingsunder logging_only (L06), where #37464 sent its raw dump and this PR sends nothing, matchingmainType
🆕 New Feature
🐛 Bug Fix
Caveats (if any)
Medium
/mcp-rest/tools/callrunspre_mcp_callguardrails twice (proxy pre-call hook with the raw body, then the MCP manager with the resolved tool name), so Singulr sees twomcp_requestscans per call with differenttool_namespellings; generic to every guardrail, unchanged hereapply_guardrailwith no proxy logging object is still classified by themcp_tool_namekey (feat(guardrails): update Singulr guardrail api contract and add logging, pre-MCP, and post-MCP hooks #37464 behaviour); every proxy route seeds the logger before the pre-call hook, so no such caller exists in the tree todayLow
tool_resultcontract is text; generic, flagged by veria as lowblock_on_error: trueare recorded asguardrail_intervenedrather thanguardrail_failed_to_respond, and a Singulr timeout surfaces as 408 regardless ofblock_on_error; both pre-existing onmain/v1/completionsunderlogging_onlyscans the response only (the logging scratch request has noprompt); pre-existing generic/v1/embeddingsunderlogging_onlyis not scanned by anyapply_guardrailguardrail: the base logging scan hands the embeddings translation user-role message dicts and it skips with a warning; pre-existing generic, left untouched here becauseinspect_embeddingsis off by default for the vendors that honour it, follow-upFinal Attestation
ran /live-pr-risk on the final tip 1460856 and found no regressions/backward incompatible risks