Skip to content

feat(proxy): add POST /management/v1/users/bulk_delete and POST /management/v1/teams/{team_id}/members/bulk_delete - #41039

Merged
ryan-crabbe-berri merged 14 commits into
mainfrom
litellm_bulk_user_delete
Sep 15, 2026
Merged

ryan-crabbe-berri merged 14 commits into
mainfrom
litellm_bulk_user_delete

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Offboarding many users means one POST /user/delete per user
  • Nothing removes a set of users from one team in a single call
  • No per-row report of what got deleted and where

How it solves it:

  • New POST /management/v1/users/bulk_delete takes up to 500 user ids in one body
  • Every deleted user is taken off every team they were on
  • New POST /management/v1/teams/{team_id}/members/bulk_delete removes up to 500 members from one team
  • Each team is rewritten once, under the team lock, from a fresh roster
  • Both return {"data": [...]} with one result per row, in request order
  • Errors are application/problem+json with urn:litellm:error:* types: 422 unknown body field, 400 unknown query param, 403 not allowed, 404 unknown team
  • Deleted keys are evicted from the auth cache on every worker, so they stop working at once, not at TTL expiry
  • Roster matching is by user_id; an entry is matched by email only when it carries no user_id (rosters written by older releases), so a teammate who shares the deleted user's email keeps their seat and keys. On the team route an id-only row is widened to that user's email only when the user's own teams lists the team, so it clears the user's email-only entry (a Bugbot finding) without touching an entry that belongs to a same-email user when the requested one is not a member (a Greptile finding)
  • One transaction per user bulk delete: team rewrites (locked in sorted team id order), then user rows, keys, invitation links and memberships. A failure anywhere rolls the whole batch back and every row reports success: false

Merging main after #41028 put POST /management/v1/users/bulk and POST /management/v1/users/bulk_delete in the same management_v1/users.py, and moved the filter that drops a parent too_short error caused by a rejected item into the shared request_validation_problem, so every /management/v1 route reports only the item's own error

Both routes follow the Management API modernization one-pager (/management/v1, plural resources, action sub-path, {data} envelope, RFC 9457 problems). There are no legacy /user/bulk_delete or /team/bulk_member_delete routes to keep: neither ever shipped, they only existed in earlier revisions of this PR

User Flow

Before: a proxy admin offboarding a batch of users has to loop over the single-user endpoints from their own script

  1. They send POST https://litellm-domain/management/v1/users/bulk_delete with {"user_ids": ["user-1", "user-2"]}
  2. They get back {"detail": "Not Found"} with HTTP 404
  3. They send POST https://litellm-domain/management/v1/teams/team-1/members/bulk_delete with {"members": [{"user_id": "user-1"}]}
  4. They get back {"detail": "Not Found"} with HTTP 404
  5. They fall back to one POST https://litellm-domain/user/delete per user and one POST https://litellm-domain/team/member_delete per member, handling failures row by row

After: the same admin sends one request per job and gets one report back

  1. They send POST https://litellm-domain/management/v1/users/bulk_delete with {"user_ids": ["user-1", "user-2", "ghost", "user-1"]}
  2. They get HTTP 200 with data holding one entry per id in the order they sent them, each with user_id, user_email, success, teams_removed and error
  3. user-1 and user-2 come back success: true with the team ids they were removed from; ghost says User id=ghost not found and the repeated user-1 says Duplicate user_id in request, both success: false
  4. GET https://litellm-domain/team/info?team_id=team-1 no longer lists them, and GET https://litellm-domain/user/info?user_id=user-1 returns HTTP 404
  5. They send POST https://litellm-domain/management/v1/teams/team-1/members/bulk_delete with {"members": [{"user_id": "user-3"}, {"user_email": "[email protected]"}, {"user_id": "nobody"}]}
  6. They get HTTP 200 with data holding one result per member; nobody says User not found in team and the other two are removed; a repeated row says Duplicate member in request
  7. Sending 501 ids, an unknown body field, or a member row carrying both user_id and user_email returns HTTP 422 application/problem+json naming the field; an unknown query parameter returns HTTP 400; an unknown team id returns HTTP 404 urn:litellm:error:team-not-found
  8. A key belonging to a deleted user is refused with HTTP 401 on the very next request on any proxy instance, even when it authenticated a moment earlier
  9. A proxy_admin_viewer or internal_user key gets HTTP 403 urn:litellm:error:forbidden on the users route; an internal_user who is not admin of the team gets the same 403 problem on the team route

Relevant issues

Affected release

Linear ticket

Resolves LIT-5958

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

QA run at 6ea1085bc3 (after, the PR tip) against e766277846 (before, the merge base before main was merged in). The after leg is two separate proxy processes, each --num_workers 2, sharing one PostgreSQL and one Redis (general_settings.coordination_redis, which is what starts the auth-cache invalidation subscriber in every worker and gives the eviction broadcast somewhere to go). Requests alternate between instance A (:41011) and instance B (:41012). Setup: /team/new for T1 and T2, /user/new for u1, u2, u3 on both teams, keep on T1, twin on T2 (then /user/update gives twin the same email as u1), legacy on T2, T3 with legacy2 and keep, outsider (no teams, same email as legacy2), /key/generate for u3, and an internal_user with its own key. Master key sk-1234, H is the master-key auth and JSON content-type headers

One fixture step is SQL rather than HTTP: today's /team/member_add always fills user_id, but rosters written by older releases can hold email-only entries, so the script strips user_id from legacy's T2 entry and legacy2's T3 entry with update "LiteLLM_TeamTable" set members_with_roles = ... before the run. /team/info on T3 then shows [{"user_id":"default_user_id","user_email":null},{"user_id":null,"user_email":"[email protected]"},{"user_id":"qa-tip4-keep","user_email":"[email protected]"}]

Before (e766277, instance on :41021)

  1. curl -s -X POST http://localhost:41021/management/v1/users/bulk_delete "${H[@]}" -d '{}' -w ' HTTP %{http_code}\n'
  2. {"detail":"Not Found"} HTTP 404
  3. curl -s -X POST http://localhost:41021/management/v1/teams/some-team/members/bulk_delete "${H[@]}" -d '{}' -w ' HTTP %{http_code}\n'
  4. {"detail":"Not Found"} HTTP 404

After (6ea1085, instances on :41011 and :41012)

Warm u3's key on both instances

  1. curl -s http://localhost:41011/models -H "Authorization: Bearer $U3K" -o /dev/null -w 'A: HTTP %{http_code}\n' gives A: HTTP 200
  2. curl -s http://localhost:41012/models -H "Authorization: Bearer $U3K" -o /dev/null -w 'B: HTTP %{http_code}\n' gives B: HTTP 200

Team members bulk_delete on instance B: by id, by email, plus one stranger

  1. curl -s http://localhost:41012/management/v1/teams/$T1/members/bulk_delete "${H[@]}" -w ' HTTP %{http_code}\n' -d '{"members":[{"user_id":"qa-tip4-u1"},{"user_email":"[email protected]"},{"user_id":"nobody-tip4"}]}'
  2. {"data":[{"user_id":"qa-tip4-u1","user_email":null,"success":true,"error":null},{"user_id":null,"user_email":"[email protected]","success":true,"error":null},{"user_id":"nobody-tip4","user_email":null,"success":false,"error":"User not found in team"}]} HTTP 200
  3. curl -s "http://localhost:41011/team/info?team_id=$T1" -H 'Authorization: Bearer sk-1234' | jq -c '[.team_info.members_with_roles[].user_id]' (instance A) gives ["default_user_id","qa-tip4-u3","qa-tip4-keep"]
  4. curl -s "http://localhost:41011/user/info?user_id=qa-tip4-u1" -H 'Authorization: Bearer sk-1234' | jq -c '.user_info.teams' gives ["e2cabf5b-ff9f-408c-9d24-0dd5b7b0f743"] (only T2 left)

A member row naming both user_id and user_email is a 422 problem

  1. curl -s http://localhost:41011/management/v1/teams/$T1/members/bulk_delete "${H[@]}" -w ' HTTP %{http_code}\n' -d '{"members":[{"user_id":"qa-tip4-keep","user_email":"[email protected]"}]}'
  2. {"type":"urn:litellm:error:invalid-request-body","title":"Invalid request body","status":422,"detail":"members.0: Value error, Each member must be identified by exactly one of user_id or user_email"} HTTP 422

Unknown team is 404, unknown body field is 422, unknown query param is 400

  1. curl -s http://localhost:41011/management/v1/teams/no-such-tip4/members/bulk_delete "${H[@]}" -w ' HTTP %{http_code} %{content_type}\n' -d '{"members":[{"user_id":"x"}]}'
  2. {"type":"urn:litellm:error:team-not-found","title":"Team not found","status":404,"detail":"Team id=no-such-tip4 does not exist in db"} HTTP 404 application/problem+json
  3. curl -s http://localhost:41011/management/v1/teams/$T1/members/bulk_delete "${H[@]}" -w ' HTTP %{http_code} %{content_type}\n' -d '{"team_id":"'$T1'","members":[{"user_id":"x"}]}'
  4. {"type":"urn:litellm:error:invalid-request-body","title":"Invalid request body","status":422,"detail":"team_id: Extra inputs are not permitted"} HTTP 422 application/problem+json
  5. curl -s 'http://localhost:41011/management/v1/users/bulk_delete?dry_run=1' "${H[@]}" -w ' HTTP %{http_code} %{content_type}\n' -d '{"user_ids":["x"]}'
  6. {"type":"urn:litellm:error:unknown-query-parameter","title":"Unknown query parameter","status":400,"detail":"Unrecognized query parameter(s): dry_run.","allowed":[]} HTTP 400 application/problem+json

An id-only row for a same-email user who is not on T3 leaves the email-only entry alone

  1. curl -s http://localhost:41012/management/v1/teams/$T3/members/bulk_delete "${H[@]}" -w ' HTTP %{http_code}\n' -d '{"members":[{"user_id":"qa-tip4-outsider"}]}'
  2. {"data":[{"user_id":"qa-tip4-outsider","user_email":null,"success":false,"error":"User not found in team"}]} HTTP 200
  3. curl -s "http://localhost:41011/team/info?team_id=$T3" -H 'Authorization: Bearer sk-1234' | jq -c '[.team_info.members_with_roles[]|{user_id,user_email}]' still gives [{"user_id":"default_user_id","user_email":null},{"user_id":null,"user_email":"[email protected]"},{"user_id":"qa-tip4-keep","user_email":"[email protected]"}]

An id-only row clears the member's own email-only roster entry (T3), read back through instance B

  1. curl -s http://localhost:41011/management/v1/teams/$T3/members/bulk_delete "${H[@]}" -w ' HTTP %{http_code}\n' -d '{"members":[{"user_id":"qa-tip4-legacy2"}]}'
  2. {"data":[{"user_id":"qa-tip4-legacy2","user_email":null,"success":true,"error":null}]} HTTP 200
  3. curl -s "http://localhost:41012/team/info?team_id=$T3" -H 'Authorization: Bearer sk-1234' | jq -c '[.team_info.members_with_roles[]|{user_id,user_email}]' gives [{"user_id":"default_user_id","user_email":null},{"user_id":"qa-tip4-keep","user_email":"[email protected]"}]
  4. curl -s "http://localhost:41012/user/info?user_id=qa-tip4-legacy2" -H 'Authorization: Bearer sk-1234' | jq -c '.user_info.teams' gives []

Users bulk_delete on instance A: three users, one ghost, one duplicate

  1. curl -s http://localhost:41011/management/v1/users/bulk_delete "${H[@]}" -w ' HTTP %{http_code}\n' -d '{"user_ids":["qa-tip4-u1","qa-tip4-u3","ghost-tip4","qa-tip4-u1","qa-tip4-legacy"]}'
  2. {"data":[{"user_id":"qa-tip4-u1","user_email":"[email protected]","success":true,"teams_removed":["e2cabf5b-ff9f-408c-9d24-0dd5b7b0f743"],"error":null},{"user_id":"qa-tip4-u3","user_email":"[email protected]","success":true,"teams_removed":["ac0475a2-1a44-4a47-9463-97334dc11749","e2cabf5b-ff9f-408c-9d24-0dd5b7b0f743"],"error":null},{"user_id":"ghost-tip4","user_email":null,"success":false,"teams_removed":[],"error":"User id=ghost-tip4 not found"},{"user_id":"qa-tip4-u1","user_email":null,"success":false,"teams_removed":[],"error":"Duplicate user_id in request: qa-tip4-u1"},{"user_id":"qa-tip4-legacy","user_email":"[email protected]","success":true,"teams_removed":["e2cabf5b-ff9f-408c-9d24-0dd5b7b0f743"],"error":null}]} HTTP 200

u3's key dies on every worker of both instances

  1. curl -s http://localhost:41012/models -H "Authorization: Bearer $U3K" -w ' B: HTTP %{http_code}\n' (instance B, which did not handle the delete) gives {"error":{"message":"Authentication Error, Invalid proxy server token passed. ... Unable to find token in cache or LiteLLM_VerificationTokenTable","type":"token_not_found_in_db","param":"key","code":"401"}} B: HTTP 401
  2. curl -s http://localhost:41011/models -H "Authorization: Bearer $U3K" -w ' A: HTTP %{http_code}\n' gives the same body and A: HTTP 401
  3. for i in 1 2 3 4; do for p in 41011 41012; do curl -s -o /dev/null -w "$p: %{http_code} " http://localhost:$p/models -H "Authorization: Bearer $U3K"; done; done gives 41011: 401 41012: 401 41011: 401 41012: 401 41011: 401 41012: 401 41011: 401 41012: 401
  4. docker exec qa-redis redis-cli pubsub numsub litellm_proxy.auth_cache_invalidation gives 4 while the proxies are up: one subscriber per worker. Earlier runs on this PR configured Redis through router_settings only, which leaves that channel with zero subscribers, and a warm worker kept answering 200 for the deleted key until its TTL. The 5ace6fa731 run's "one 200" was that misconfiguration, not delivery lag

Rosters and lookups read back through instance B

  1. curl -s "http://localhost:41012/team/info?team_id=$T2" -H 'Authorization: Bearer sk-1234' | jq -c '[.team_info.members_with_roles[]|{user_id,user_email}]' gives [{"user_id":"default_user_id","user_email":null},{"user_id":"qa-tip4-u2","user_email":"[email protected]"},{"user_id":"qa-tip4-twin","user_email":"[email protected]"}]: the email-only legacy entry is gone, twin (whose user row carries u1's email) stays
  2. curl -s "http://localhost:41012/team/info?team_id=$T1" -H 'Authorization: Bearer sk-1234' | jq -c '[.team_info.members_with_roles[].user_id]' gives ["default_user_id","qa-tip4-keep"]
  3. curl -s "http://localhost:41012/user/info?user_id=qa-tip4-u3" -H 'Authorization: Bearer sk-1234' -w ' HTTP %{http_code}\n' gives {"error":{"message":"User qa-tip4-u3 not found","type":"internal_server_error","param":null,"code":"404"}} HTTP 404
  4. curl -s "http://localhost:41012/user/info?user_id=qa-tip4-twin" -H 'Authorization: Bearer sk-1234' | jq -c '{user_id: .user_id, teams: .user_info.teams}' gives {"user_id":"qa-tip4-twin","teams":["e2cabf5b-ff9f-408c-9d24-0dd5b7b0f743"]}
  5. curl -s "http://localhost:41012/key/list?user_id=qa-tip4-u3&return_full_object=true" -H 'Authorization: Bearer sk-1234' | jq -c '{total_count}' gives {"total_count":0}

An internal_user key is refused with a 403 problem

  1. curl -s http://localhost:41011/management/v1/users/bulk_delete -H "Authorization: Bearer $INTERNAL_USER_KEY" -H 'Content-Type: application/json' -w ' HTTP %{http_code} %{content_type}\n' -d '{"user_ids":["qa-tip4-u2"]}'
  2. {"type":"urn:litellm:error:forbidden","title":"Forbidden","status":403,"detail":"Only PROXY_ADMIN or ORG_ADMIN users may delete users."} HTTP 403 application/problem+json
  3. curl -s http://localhost:41012/management/v1/teams/$T2/members/bulk_delete -H "Authorization: Bearer $INTERNAL_USER_KEY" -H 'Content-Type: application/json' -w ' HTTP %{http_code} %{content_type}\n' -d '{"members":[{"user_id":"qa-tip4-u2"}]}'
  4. {"type":"urn:litellm:error:forbidden","title":"Forbidden","status":403,"detail":"Call not allowed. User not proxy admin OR team admin OR org admin for this team. route='/management/v1/teams/e2cabf5b-ff9f-408c-9d24-0dd5b7b0f743/members/bulk_delete'"} HTTP 403 application/problem+json

501 ids is a 422 problem

  1. python3 -c 'import json;print(json.dumps({"user_ids":[f"x{i}" for i in range(501)]}))' > /tmp/ids501.json && curl -s http://localhost:41011/management/v1/users/bulk_delete "${H[@]}" -d @/tmp/ids501.json -w ' HTTP %{http_code} %{content_type}\n'
  2. {"type":"urn:litellm:error:invalid-request-body","title":"Invalid request body","status":422,"detail":"user_ids: Tuple should have at most 500 items after validation, not 501"} HTTP 422 application/problem+json

Observations from the run, none caused or changed by this PR:

  • Existing /user/delete rejects the whole batch on one unknown id
  • The roster keeps the email a member had when added; /user/update changing twin's email does not rewrite T2's entry

Dependent paths this PR touches (e766277846 vs 6ea1085bc3)

The run above was first taken at ca04b03c2c, repeated at b0acda2825 (type annotations only), at 5ace6fa731 (the roster matching fix, with the T3 step added) at 98ea1475 (the same-email guard, with the outsider step added and Redis wired as coordination_redis) and at 6ea1085bc3 (the merge of main after #41028 landed); every earlier step answered the same each time, and the 6ea1085bc3 run diffs clean against the 98ea1475 one once keys and ids are masked. Three dependents outside the new routes changed under them, so each was driven on the base instance (:41021, e766277846) and the tip instance (:41012) with the same commands, and the outputs diff clean once ids and timestamps are masked. check_route_access now backs the view-only blocklists, so a proxy_admin_viewer key hit /user/new, /team/new, /key/{id}/regenerate (all HTTP 403, same message) and /user/info, /team/list (both HTTP 200). PrismaClient.tx() grew a timeout parameter whose default is Prisma's own 5s, so /team/member_delete, /user/delete, /team/delete were run on a fresh team and member (all HTTP 200, same bodies). The /management/v1 validation handler now splits body and query errors, so the existing list routes were hit with /management/v1/budgets?page_size=abc, ?page_size=0, ?page=-1, ?foo=1 (all HTTP 400 application/problem+json with the same invalid-query-parameter / unknown-query-parameter type, title and detail on both sides). Every entry in both blocklists was also checked offline against check_route_access and plain in for itself and for prefixed, suffixed, upper-cased and truncated variants, with no divergence; no pre-existing entry carries a placeholder or a regex metacharacter, so the pattern branch only ever fires for the new {team_id} route

  1. curl -s http://localhost:41021/management/v1/budgets?page_size=abc -H "Authorization: Bearer $MASTER_KEY" -w '\nHTTP %{http_code} %{content_type}\n'
  2. {"type":"urn:litellm:error:invalid-query-parameter","title":"Invalid query parameter","status":400,"detail":"'page_size' must be an integer."} then HTTP 400 application/problem+json (identical on :41012)
  3. curl -s -X POST http://localhost:41021/team/new -H "Authorization: Bearer $VIEWER_KEY" -H 'Content-Type: application/json' -d '{"team_alias":"lpr-x"}' -w '\nHTTP %{http_code}\n'
  4. {"error":{"message":"user not allowed to access this route, role= proxy_admin_viewer. Trying to access: /team/new","type":"auth_error","param":"None","code":"403"}} then HTTP 403 (identical on :41012)

Type

🆕 New Feature

Caveats (if any)

Severe

  • Without Redis, a deleted user's key keeps working on other workers and instances until its auth cache TTL expires. That is how every management mutation behaves today (the broadcast is a no-op with no Redis), and the run above shows it working once Redis is configured as general_settings.coordination_redis (or the REDIS_* env fallback). A router_settings or cache Redis alone does not start the invalidation subscriber

Medium

  • Users bulk_delete holds one transaction, and one advisory lock per affected team, for the whole batch. The transaction timeout is 60s; a batch whose users span hundreds of teams rewrites them serially inside it, and a timeout rolls back every row of that request (each reported success: false, the caller retries with a smaller batch)
  • route_checks.py now matches management routes by pattern (so /management/v1/teams/{team_id}/... entries match a concrete team id). The existing tests in tests/test_litellm/proxy/auth/test_route_checks.py pass, but this widens matching for every parametrized entry in those lists, not only the new ones

Low

  • The team route creates no per-member audit log entry, matching /team/member_delete
  • Org admins are checked the same way as /user/delete: every organization the target belongs to must be one they administer, and users with no organization are out of scope
  • The bulk shape (.../bulk_delete action) is the one place the one-pager leaves open ("bulk operations deferred"); it is a declared action on the collection, which is what the doc asks for non-CRUD behavior
  • CI Vertex AI / Run tests fails on two embedding cost assertions in test_batch_embed_content_transformation.py. This branch touches no Vertex code: the same two tests fail the same way at the merge base e766277846 with the published price map and pass with LITELLM_LOCAL_MODEL_COST_MAP=True, so the remote price sync for gemini-embedding moved the numbers under them
  • Analyze (python) (CodeQL) has failed on every run of this branch with Result set is larger than the limit of 2GiB, and recent main runs hit the same evaluator limit

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR
  • 6ea1085 passes /live-pr-risk

Link to Devin session: https://app.devin.ai/sessions/819748a6114e4e329be11d97a69f301c
Open in Devin Desktop: https://app.devin.ai/desktop/session/819748a6114e4e329be11d97a69f301c?variant=devin

ryan-crabbe-berri and others added 2 commits September 14, 2026 04:55
…lete

Batch user deletion that also removes each user from every team they belong to, and batch removal of many members from one team. Each touched team is rewritten once under the team advisory lock from a roster re-read under that lock

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…opping them

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration
devin-ai-integration Bot requested a review from a team September 14, 2026 04:58
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@codspeed

codspeed Bot commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_bulk_user_delete (6ea1085) with main (08b4332)

Open in CodSpeed

@greptile-apps

greptile-apps Bot commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR adds bulk user deletion and bulk team-member removal to the Management v1 API.

  • Adds ordered, per-row outcomes for batches of up to 500 entries.
  • Keeps user, membership, key, invitation, and team-roster changes transactionally aligned.
  • Serializes team roster rewrites with advisory locks and fresh roster reads.
  • Evicts deleted credentials locally and broadcasts invalidation to other workers.
  • Adds Management API validation problems, authorization checks, generated schemas, and focused behavior tests.

Confidence Score: 5/5

The PR appears safe to merge; no outstanding previous finding or actionable new defect remains.

The current implementation validates exclusive member identifiers, limits cleanup to matched users, performs bulk user deletion atomically, disambiguates legacy email-only roster entries using membership state, and invalidates deleted credentials across configured workers. Every previous thread was manually resolved, and the changes since the prior review did not introduce a new route or validation regression.

Important Files Changed

Filename Overview
litellm/proxy/management_helpers/bulk_user_deletion.py Implements transactional bulk deletion, locked roster reconciliation, authorization, audit handling, and cross-worker credential invalidation.
litellm/proxy/management_endpoints/management_v1/users.py Registers the bulk user creation and deletion actions with separate schemas and Management API problem handling.
litellm/proxy/management_endpoints/management_v1/teams.py Exposes the authenticated bulk team-member deletion action.
litellm/proxy/list_api/common.py Centralizes Management API validation errors and removes redundant collection-length errors caused by rejected nested items.
litellm/proxy/auth/route_checks.py Applies existing viewer write restrictions to the new concrete and parameterized management routes.
litellm/proxy/utils.py Allows callers to select a bounded Prisma transaction timeout while preserving the existing five-second default.

Reviews (10): Last reviewed commit: "Merge branch 'main' into litellm_bulk_us..." | Re-trigger Greptile

Comment thread litellm/proxy/management_helpers/bulk_user_deletion.py Outdated
Comment thread litellm/types/proxy/management_endpoints/team_endpoints.py Outdated
Comment thread litellm/proxy/management_helpers/bulk_user_deletion.py Outdated
Comment thread litellm/proxy/management_helpers/bulk_user_deletion.py Outdated
Comment thread litellm/proxy/management_helpers/bulk_user_deletion.py Outdated
Comment thread litellm/proxy/management_helpers/bulk_user_deletion.py Outdated
Comment thread litellm/proxy/management_endpoints/internal_user_endpoints.py Outdated
@codecov

codecov Bot commented Sep 14, 2026 •

Copy link
Copy Markdown

… deletion transactional

/user/bulk_delete now deletes the users' keys, invitation links, org and team
memberships and user rows in one transaction and reports a rolled-back batch
per row instead of leaving partial deletes behind. Both bulk endpoints evict
the deleted keys (and deleted user objects) from the auth cache, so a deleted
key stops authenticating immediately rather than at TTL expiry.

/team/bulk_member_delete rejects member rows that carry both user_id and
user_email, reports repeated rows as duplicates, and only cleans up keys and
memberships of members it actually matched.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

…ne transaction

Lock affected teams in sorted order inside a single 60s transaction so a
failure on any team rolls back every rewrite and every user row delete.
PrismaClient.tx() gains an optional timeout for the larger batch.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/proxy/management_helpers/bulk_user_deletion.py
Comment thread litellm/proxy/management_helpers/bulk_user_deletion.py Outdated
ryan-crabbe-berri and others added 2 commits September 14, 2026 05:56
…r route coverage

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…and give /team/bulk_member_delete the 60s batch timeout

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread litellm/proxy/management_helpers/bulk_user_deletion.py
…thout touching same-email teammates

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

ryan-crabbe-berri and others added 2 commits September 14, 2026 22:00
…gement/v1

Replaces POST /user/bulk_delete and POST /team/bulk_member_delete with
POST /management/v1/users/bulk_delete and
POST /management/v1/teams/{team_id}/members/bulk_delete per the Management
API modernization one-pager: {data} envelopes, application/problem+json
errors with urn:litellm:error:* types, 422 on unknown body fields, 400 on
unknown query params, 403 on authorization failures, 404 on unknown team.

Route checks now match parametrized management/v1 paths so team-scoped
callers reach the endpoint's own authorization and get a 403 problem
instead of the generic 401.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration devin-ai-integration Bot changed the title feat(proxy): add POST /user/bulk_delete and POST /team/bulk_member_delete feat(proxy): add POST /management/v1/users/bulk_delete and POST /management/v1/teams/{team_id}/members/bulk_delete Sep 14, 2026
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/proxy/management_helpers/bulk_user_deletion.py
…l-only roster entry

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread litellm/proxy/management_helpers/bulk_user_deletion.py
…n the user lists the team

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 98ea147. Configure here.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@ryan-crabbe-berri
ryan-crabbe-berri merged commit 8b6c398 into main Sep 15, 2026
91 of 92 checks passed
@ryan-crabbe-berri
ryan-crabbe-berri deleted the litellm_bulk_user_delete branch September 15, 2026 18:12
doonga pushed a commit to greyrock-labs/home-ops that referenced this pull request Sep 28, 2026
…103.0) (#290)

This PR contains the following updates:

| Package | Update | Change |
|---|---|---|
| [ghcr.io/berriai/litellm](https://images.chainguard.dev/directory/image/wolfi-base/overview) ([source](https://github.com/BerriAI/litellm)) | minor | `v1.102.1` → `v1.103.0` |

---

### Release Notes

<details>
<summary>BerriAI/litellm (ghcr.io/berriai/litellm)</summary>

### [`v1.103.0`](https://github.com/BerriAI/litellm/releases/tag/v1.103.0)

[Compare Source](https://github.com/BerriAI/litellm/compare/v1.102.1...v1.103.0)

#### Verify Docker Image Signature

All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0).

**Verify using the pinned commit hash (recommended):**

A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \
  ghcr.io/berriai/litellm:v1.103.0
```

**Verify using the release tag (convenience):**

Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/v1.103.0/cosign.pub \
  ghcr.io/berriai/litellm:v1.103.0
```

Expected output:

```
The following checks were performed on each of these signatures:
  - The cosign claims were validated
  - The signatures were verified against the specified public key
```

***

#### What's Changed

- fix(responses): translate the reasoning object into a chat-completion reasoning effort by [@&#8203;joshgarnett](https://github.com/joshgarnett) in [#&#8203;36363](https://github.com/BerriAI/litellm/pull/36363)
- fix(proxy): bound tool and guardrail index create\_many by the spend-log statement budgets by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40561](https://github.com/BerriAI/litellm/pull/40561)
- fix(mcp): require admission for delegated OAuth by [@&#8203;joshua-berri](https://github.com/joshua-berri) in [#&#8203;40923](https://github.com/BerriAI/litellm/pull/40923)
- fix(logging): log one bounded summary for a burst of timed-out LoggingWorker callbacks by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40912](https://github.com/BerriAI/litellm/pull/40912)
- fix(fireworks): resolve short model names to long cost map keys by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40929](https://github.com/BerriAI/litellm/pull/40929)
- ci: remove main branch source guard by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;40172](https://github.com/BerriAI/litellm/pull/40172)
- chore(ci): promote internal staging to main by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;40942](https://github.com/BerriAI/litellm/pull/40942)
- fix(spend\_logs): store litellm\_call\_id and match it in request\_id lookups by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;39068](https://github.com/BerriAI/litellm/pull/39068)
- fix(auth): refresh lite login session token grants from the live user and team rows by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;40657](https://github.com/BerriAI/litellm/pull/40657)
- feat(bedrock): support file delete and list for S3-backed managed files by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;39836](https://github.com/BerriAI/litellm/pull/39836)
- fix(proxy): gate the webhook test alert on proxy admins by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;40814](https://github.com/BerriAI/litellm/pull/40814)
- fix(ui): hide admin write-form tabs on the models page from view-only admins by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;38867](https://github.com/BerriAI/litellm/pull/38867)
- fix(anthropic-adapter): surface mid-stream provider errors as Anthropic error events by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33352](https://github.com/BerriAI/litellm/pull/33352)
- docs(github): add an Affected release section to the PR template by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;40618](https://github.com/BerriAI/litellm/pull/40618)
- docs(e2e): ban unit tests under tests/e2e by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33852](https://github.com/BerriAI/litellm/pull/33852)
- fix(router): preserve Azure Entra ID params in reusable credentials by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40889](https://github.com/BerriAI/litellm/pull/40889)
- docs(user endpoints): remove unsupported soft\_budget param from user docstrings by [@&#8203;shivamrawat1](https://github.com/shivamrawat1) in [#&#8203;36585](https://github.com/BerriAI/litellm/pull/36585)
- feat(friendli): auto-sync Friendli model metadata into price registry by [@&#8203;Lee-Si-Yoon](https://github.com/Lee-Si-Yoon) in [#&#8203;35918](https://github.com/BerriAI/litellm/pull/35918)
- build(deps): bump smol-toml to 1.8.0 to clear GHSA-7w5x-hrqm-74c2 in osv-scan by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40478](https://github.com/BerriAI/litellm/pull/40478)
- fix(bedrock\_mantle): price GovCloud regions from the regional cost row and accept region-prefixed model names by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;39846](https://github.com/BerriAI/litellm/pull/39846)
- chore(ci): remerge internal staging by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;40943](https://github.com/BerriAI/litellm/pull/40943)
- test(auth): freeze the cache clock in auth prefetch tests by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40996](https://github.com/BerriAI/litellm/pull/40996)
- feat(jwt): allow virtual\_key\_claim\_field per issuer by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40927](https://github.com/BerriAI/litellm/pull/40927)
- fix(cost): bill cached realtime audio tokens at the audio cache-read rate by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40627](https://github.com/BerriAI/litellm/pull/40627)
- perf(logging): skip correlation contextvar stamping when request\_correlation\_in\_logs is off by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41054](https://github.com/BerriAI/litellm/pull/41054)
- feat(pricing): add azure gpt-chat-latest rates and drop retired friendliai llama-3.1 entries by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40976](https://github.com/BerriAI/litellm/pull/40976)
- build(deps): re-suppress GHSA-h7x2-h6g9-p789 in osv-scan on main, mlflow still has no fixed release by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41104](https://github.com/BerriAI/litellm/pull/41104)
- fix(otel): cap per-index OpenInference message attributes span-wide by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40562](https://github.com/BerriAI/litellm/pull/40562)
- feat(proxy): add general\_settings.allowed\_file\_extensions for /v1/files uploads by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41106](https://github.com/BerriAI/litellm/pull/41106)
- fix(proxy): forward provider request id headers on mapped error responses by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40925](https://github.com/BerriAI/litellm/pull/40925)
- fix(router): name the all-deployments-in-cooldown error on 429 responses by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40995](https://github.com/BerriAI/litellm/pull/40995)
- fix(ui): show the team alias on the model info page and in its raw JSON by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40992](https://github.com/BerriAI/litellm/pull/40992)
- feat(proxy): honor LITELLM\_DISABLE\_ACCESS\_LOG\_PATHS to drop noisy uvicorn access log lines by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41096](https://github.com/BerriAI/litellm/pull/41096)
- fix(prometheus): label pre-call rate limit failures with the resolved api\_provider by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41059](https://github.com/BerriAI/litellm/pull/41059)
- perf(proxy): serialize /model/info listing once with orjson by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41114](https://github.com/BerriAI/litellm/pull/41114)
- fix(utils): stop wrapper\_async submitting the sync success handler twice by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41115](https://github.com/BerriAI/litellm/pull/41115)
- fix(redis): log a timeout streak once per interval instead of one line per cache call by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40817](https://github.com/BerriAI/litellm/pull/40817)
- fix(router): record flat retry attempts and cap retries from attempted\_retries by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40930](https://github.com/BerriAI/litellm/pull/40930)
- refactor(prometheus): source PROXY\_LLM\_PROVIDER\_FALLBACK from litellm.constants by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41118](https://github.com/BerriAI/litellm/pull/41118)
- fix(proxy): hide default credentials login hint when UI\_PASSWORD is set by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41107](https://github.com/BerriAI/litellm/pull/41107)
- fix(cli): show routed models and session stats for LLM API keys by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41116](https://github.com/BerriAI/litellm/pull/41116)
- fix(bedrock/realtime): propagate deferred Nova Sonic stream failures to the router by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41064](https://github.com/BerriAI/litellm/pull/41064)
- fix(proxy): keep org admins' own team memberships in other orgs visible on team list by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41086](https://github.com/BerriAI/litellm/pull/41086)
- feat(model\_info): provider-scoped fill\_missing\_for\_providers backfill from fallback generalization rules by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41093](https://github.com/BerriAI/litellm/pull/41093)
- fix(auth): load team membership once per request and skip prisma on an L1 hit by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41102](https://github.com/BerriAI/litellm/pull/41102)
- refactor(harness): expand independent trace coverage by [@&#8203;yujonglee-berri](https://github.com/yujonglee-berri) in [#&#8203;41120](https://github.com/BerriAI/litellm/pull/41120)
- fix(proxy): release max\_parallel\_requests slot when a realtime session ends without LLM callbacks by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41113](https://github.com/BerriAI/litellm/pull/41113)
- fix(router): cool down team deployments on 429 when a sibling serves the same public model by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40991](https://github.com/BerriAI/litellm/pull/40991)
- fix(ui): move tags typed into key metadata JSON into the Tags field by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41023](https://github.com/BerriAI/litellm/pull/41023)
- fix(ui): let team admins grant a team all proxy models by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;40196](https://github.com/BerriAI/litellm/pull/40196)
- chore(lint): graduate 12 rules from the strict-gate ratchet by [@&#8203;HUAHAODIA](https://github.com/HUAHAODIA) in [#&#8203;41048](https://github.com/BerriAI/litellm/pull/41048)
- test: add dedicated CircleCI integration contract foundation by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41066](https://github.com/BerriAI/litellm/pull/41066)
- fix(utils): keep litellm params out of provider request bodies by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41018](https://github.com/BerriAI/litellm/pull/41018)
- fix(openai): keep extra\_headers out of the chat request body on the httpx handler path by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41141](https://github.com/BerriAI/litellm/pull/41141)
- test: cover persisted updates and warmed authorization policies by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41070](https://github.com/BerriAI/litellm/pull/41070)
- fix(proxy): resolve x-litellm-call-id from response metadata when routes omit call\_id by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41056](https://github.com/BerriAI/litellm/pull/41056)
- chore(prices): sync Vertex AI prices: 14 models by [@&#8203;berriai-litellm-provider-info-sync](https://github.com/berriai-litellm-provider-info-sync)\[bot] in [#&#8203;40955](https://github.com/BerriAI/litellm/pull/40955)
- ci(codeql): exclude noisy Python quality queries by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41142](https://github.com/BerriAI/litellm/pull/41142)
- feat(proxy): predict prompt-cache costs across deployments by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;40877](https://github.com/BerriAI/litellm/pull/40877)
- fix(prompt\_security): keep polling file sanitization through non-terminal statuses by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41131](https://github.com/BerriAI/litellm/pull/41131)
- fix(health): skip background health check DB writes when the latest-row read fails by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41145](https://github.com/BerriAI/litellm/pull/41145)
- feat(model\_armor): logging\_only mode scans completed streams after delivery by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40702](https://github.com/BerriAI/litellm/pull/40702)
- fix(bedrock guardrails): derive contextual grounding source and query from plain messages by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41132](https://github.com/BerriAI/litellm/pull/41132)
- fix(cli): drop enum.StrEnum so the CLI imports on Python 3.10 by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41046](https://github.com/BerriAI/litellm/pull/41046)
- fix(responses): route mid-stream error events through exception\_type so content\_policy\_fallbacks fire by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40988](https://github.com/BerriAI/litellm/pull/40988)
- fix(cost): bill gemini-embedding-2 per token and stop double charging audio by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41157](https://github.com/BerriAI/litellm/pull/41157)
- test: bind management E2E callers and isolate JWT actors by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;40892](https://github.com/BerriAI/litellm/pull/40892)
- fix(headroom): protect cache\_control-marked rows anywhere in history by [@&#8203;rad-p44](https://github.com/rad-p44) in [#&#8203;40315](https://github.com/BerriAI/litellm/pull/40315)
- test: add strict stateless provider replay identity by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41149](https://github.com/BerriAI/litellm/pull/41149)
- fix(ci): test checked-out model pricing in unit jobs by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41181](https://github.com/BerriAI/litellm/pull/41181)
- fix(guardrails): write per-message guardrail rewrites back onto Responses input items by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40939](https://github.com/BerriAI/litellm/pull/40939)
- fix(proxy): log the provider usage on deferred /v1/messages calls and price cache writes without a creation rate by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41172](https://github.com/BerriAI/litellm/pull/41172)
- fix(guardrails): record not\_run evaluation when scoping leaves nothing to scan by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;39050](https://github.com/BerriAI/litellm/pull/39050)
- fix(responses): preserve provider affinity by [@&#8203;AaronHowell](https://github.com/AaronHowell) in [#&#8203;40228](https://github.com/BerriAI/litellm/pull/40228)
- fix(sdk): keep body and proxy headers on BadRequestError mapped from a litellm\_proxy 400 by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40994](https://github.com/BerriAI/litellm/pull/40994)
- fix(responses): hoist Codex additional\_tools input items into the chat bridge tools by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40989](https://github.com/BerriAI/litellm/pull/40989)
- fix(router): honor team and key provider weights by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41072](https://github.com/BerriAI/litellm/pull/41072)
- test(e2e): verify streamed answers and tool continuation by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41194](https://github.com/BerriAI/litellm/pull/41194)
- fix(cli): label router costs and simplify the routed-model header by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41186](https://github.com/BerriAI/litellm/pull/41186)
- test(spend): reconcile concurrent requests and daily activity by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41188](https://github.com/BerriAI/litellm/pull/41188)
- fix(guardrails): scan the Anthropic top-level system prompt and tool\_use arguments by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40984](https://github.com/BerriAI/litellm/pull/40984)
- fix(router): count num\_retries\_per\_request across fallback hops by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41191](https://github.com/BerriAI/litellm/pull/41191)
- fix(bedrock): grant rerank, retrieve, agent, and agentcore actions in the web identity session policy by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41168](https://github.com/BerriAI/litellm/pull/41168)
- fix(vertex-live): bill Gemini Live sessions end to end (internal copy of [#&#8203;37075](https://github.com/BerriAI/litellm/issues/37075)) by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40915](https://github.com/BerriAI/litellm/pull/40915)
- fix(health): resolve litellm\_credential\_name in realtime health checks by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41173](https://github.com/BerriAI/litellm/pull/41173)
- feat(proxy): unified custom\_key\_policy hook for key generate, update and regenerate by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40921](https://github.com/BerriAI/litellm/pull/40921)
- fix(proxy): enforce custom\_key\_update policy on /key/regenerate by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40695](https://github.com/BerriAI/litellm/pull/40695)
- fix(router): preserve session model choice within each complexity tier by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41174](https://github.com/BerriAI/litellm/pull/41174)
- test(pricing): let synced GovCloud Bedrock rows cite the AWS price list by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41263](https://github.com/BerriAI/litellm/pull/41263)
- docs(github): ask for interactive coding-tool proof in the PR template by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41257](https://github.com/BerriAI/litellm/pull/41257)
- feat(proxy): add POST /management/v1/users/bulk for batched user and team membership creation by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41028](https://github.com/BerriAI/litellm/pull/41028)
- fix(credentials): answer 409 on a credential name collision, make Terraform adoption opt-in by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;40917](https://github.com/BerriAI/litellm/pull/40917)
- feat(proxy): add POST /management/v1/users/bulk\_delete and POST /management/v1/teams/{team\_id}/members/bulk\_delete by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41039](https://github.com/BerriAI/litellm/pull/41039)
- fix(proxy): list directly assigned team models in model access errors by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41256](https://github.com/BerriAI/litellm/pull/41256)
- feat(auto-router): allow opted-in team members to manage their routers by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41175](https://github.com/BerriAI/litellm/pull/41175)
- build(rust-bridge): add typed \_native stub and validate it with mypy.stubtest by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41180](https://github.com/BerriAI/litellm/pull/41180)
- feat(guardrails): add new upstream presidio pii entities including german set by [@&#8203;MvdB](https://github.com/MvdB) in [#&#8203;36775](https://github.com/BerriAI/litellm/pull/36775)
- fix(responses): filter bridged kwargs like the native Responses path by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41144](https://github.com/BerriAI/litellm/pull/41144)
- test(e2e): cover the reliability retry, cooldown, fallback, and routing-strategy cells by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;39857](https://github.com/BerriAI/litellm/pull/39857)
- fix(anthropic): add the per-turn-control beta when a message carries output\_config by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41189](https://github.com/BerriAI/litellm/pull/41189)
- fix(router): bind per-request routing\_strategy override selectors to the request's callbacks by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41178](https://github.com/BerriAI/litellm/pull/41178)
- feat(proxy): bind JWT claims to registered agents via agent\_id\_jwt\_field by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40904](https://github.com/BerriAI/litellm/pull/40904)
- fix(proxy): enforce organization budgets when max\_budget is 0 by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;41271](https://github.com/BerriAI/litellm/pull/41271)
- fix(alerting): send llm\_exceptions Slack alert for 5xx HTTPException and ProxyException by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41125](https://github.com/BerriAI/litellm/pull/41125)
- fix(headroom): protect the cached prefix through the last cache\_control breakpoint by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41161](https://github.com/BerriAI/litellm/pull/41161)
- fix(utils): cache custom HuggingFace tokenizers across /utils/token\_counter requests by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41216](https://github.com/BerriAI/litellm/pull/41216)
- fix(router): keep weighted routing when a deployment id equals a model\_name by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41156](https://github.com/BerriAI/litellm/pull/41156)
- feat(router): add capability classifier as Fuse foundation by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41270](https://github.com/BerriAI/litellm/pull/41270)
- fix(proxy): keep access-group raw SQL writes on the writer while writer\_unavailable is stale by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41283](https://github.com/BerriAI/litellm/pull/41283)
- fix(prometheus): count 401 auth failures in litellm\_proxy\_failed\_requests\_metric by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41170](https://github.com/BerriAI/litellm/pull/41170)
- test: drop tests that pin vendor facts and add the CLAUDE.md rule by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41269](https://github.com/BerriAI/litellm/pull/41269)
- fix(proxy): run the remaining inline token counts off the event loop by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40262](https://github.com/BerriAI/litellm/pull/40262)
- fix(proxy): log blocked streaming guardrail responses as failures, not success by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40191](https://github.com/BerriAI/litellm/pull/40191)
- feat(proxy): add tpd\_limit (tokens per day) for batch submissions by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40997](https://github.com/BerriAI/litellm/pull/40997)
- fix(proxy): reconcile budget reservation before enqueuing spend to the DB by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40310](https://github.com/BerriAI/litellm/pull/40310)
- fix(xai): stop sending web\_search\_options to xAI's retired Live Search path by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;38278](https://github.com/BerriAI/litellm/pull/38278)
- feat(terraform): add tpm\_limit, rpm\_limit, budget\_duration, allowed\_models to litellm\_team\_member\_add by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;38682](https://github.com/BerriAI/litellm/pull/38682)
- fix(rerank): bill Vertex search\_units from input records and give every rerank response a unique id by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35180](https://github.com/BerriAI/litellm/pull/35180)
- fix(router): stop counting caller-set timeout 408s toward deployment cooldown by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41230](https://github.com/BerriAI/litellm/pull/41230)
- feat(router): add Fuse V2 classifier after capability forecasting by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41272](https://github.com/BerriAI/litellm/pull/41272)
- fix(proxy): keep client User-Agent on auth failure spend logs by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41291](https://github.com/BerriAI/litellm/pull/41291)
- fix(proxy): reset budgets by decrementing pre-reset spend instead of zeroing rows by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41279](https://github.com/BerriAI/litellm/pull/41279)
- fix(xai): honor nested web\_search filters on the xAI Responses API by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;38268](https://github.com/BerriAI/litellm/pull/38268)
- fix(router): stop registering a caller-supplied credential as a router deployment by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;41289](https://github.com/BerriAI/litellm/pull/41289)
- fix(router): accept custom\_provider\_map providers before the first completion call by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41300](https://github.com/BerriAI/litellm/pull/41300)
- fix(proxy): return 400 instead of 500 for lone surrogate escapes in request body by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41297](https://github.com/BerriAI/litellm/pull/41297)
- fix(langsmith): keep events appended during an in-flight flush instead of clearing them by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41288](https://github.com/BerriAI/litellm/pull/41288)
- fix(logging): track spend for streams a deployment hook converted to non-streaming by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41171](https://github.com/BerriAI/litellm/pull/41171)
- fix(bedrock): sanitize client tool\_call ids to Bedrock toolUseId constraints by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40872](https://github.com/BerriAI/litellm/pull/40872)
- fix(passthrough): attribute Vertex passthrough successes to the resolved router deployment by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41307](https://github.com/BerriAI/litellm/pull/41307)
- feat(ui): persist Models table search, filters, sort and page in the URL by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41296](https://github.com/BerriAI/litellm/pull/41296)
- fix(jwt-auth): scope JWT key mappings by issuer to prevent cross-issuer collisions by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;41281](https://github.com/BerriAI/litellm/pull/41281)
- feat(openai): add openai\_system\_messages\_first to put system messages first for prompt caching by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41304](https://github.com/BerriAI/litellm/pull/41304)
- feat(ui): add custom request headers to the API Playground by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41309](https://github.com/BerriAI/litellm/pull/41309)
- feat(cli): sync Codex /model picker from proxy /v1/models in lite codex by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40476](https://github.com/BerriAI/litellm/pull/40476)
- chore: bump litellm-enterprise 0.1.67 -> 0.1.68, litellm-proxy-extras 0.4.97 -> 0.4.98, litellm 1.102.0 -> 1.103.0 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41321](https://github.com/BerriAI/litellm/pull/41321)
- feat: add aihubmix provider pricing entries by [@&#8203;IToSSc](https://github.com/IToSSc) in [#&#8203;41179](https://github.com/BerriAI/litellm/pull/41179)
- feat(auto-router): add per-model Fast mode toggle by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41282](https://github.com/BerriAI/litellm/pull/41282)
- fix(proxy): include litellm\_call\_id in LLM API exception logs by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41205](https://github.com/BerriAI/litellm/pull/41205)
- fix(proxy): keep yaml pass-through endpoints visible to auth after db overlay by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41303](https://github.com/BerriAI/litellm/pull/41303)
- fix(proxy): resolve router\_settings.model\_group\_alias before key/team model auth by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41308](https://github.com/BerriAI/litellm/pull/41308)
- fix(ui): block usage export and flag the range when a spend page fails by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;41294](https://github.com/BerriAI/litellm/pull/41294)
- fix(vertex\_ai): bill Gemini Omni Interactions usage and Veo sampleCount on passthrough by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41322](https://github.com/BerriAI/litellm/pull/41322)
- fix(proxy): honor LITELLM\_LOG for uvicorn and proxy extras loggers by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41306](https://github.com/BerriAI/litellm/pull/41306)
- fix(cost): price native Responses WebSocket turns at their returned service\_tier by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41318](https://github.com/BerriAI/litellm/pull/41318)
- fix(proxy): key model rpm/tpm override takes precedence over team model limit by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41302](https://github.com/BerriAI/litellm/pull/41302)
- fix(proxy): track per-member organization spend by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41255](https://github.com/BerriAI/litellm/pull/41255)
- feat(proxy): add /nvidia\_nim passthrough route for NIM object detection and OCR /v1/infer by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41316](https://github.com/BerriAI/litellm/pull/41316)
- feat(model\_info): add provider-neutral Gemini 2.5+ chat baseline fallback generalization by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41320](https://github.com/BerriAI/litellm/pull/41320)
- fix(spend): sum multi-round session duration in logs UI by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35388](https://github.com/BerriAI/litellm/pull/35388)
- feat(router): limit unlicensed Capability and Fuse v2 routers to one each by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41326](https://github.com/BerriAI/litellm/pull/41326)
- fix(guardrails): resolve caller identity from metadata buckets in custom code guardrail by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41126](https://github.com/BerriAI/litellm/pull/41126)
- fix(e2e): onboard dashboard users through invitations by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41319](https://github.com/BerriAI/litellm/pull/41319)
- feat(guardrails): add Microsoft Agent 365 MCP tool-call guardrail by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;38241](https://github.com/BerriAI/litellm/pull/38241)
- test: drop remaining tests that pin cost-map vendor facts by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41298](https://github.com/BerriAI/litellm/pull/41298)
- feat(ui): show average response time per model in usage model activity by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41313](https://github.com/BerriAI/litellm/pull/41313)
- fix(proxy): preserve Anthropic pricing modifiers in router savings by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41341](https://github.com/BerriAI/litellm/pull/41341)
- feat(guardrails): support pre\_call and during\_call modes for llm\_as\_a\_judge by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41128](https://github.com/BerriAI/litellm/pull/41128)
- fix(gemini): propagate the provider's modelVersion to the response model by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41338](https://github.com/BerriAI/litellm/pull/41338)
- fix(fireworks-ai): bill cache-write, reasoning and audio tokens via the shared cost calculator by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41339](https://github.com/BerriAI/litellm/pull/41339)
- feat(guardrails): singulr v2 API contract with logging\_only, pre\_mcp\_call and post\_mcp\_call by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;41329](https://github.com/BerriAI/litellm/pull/41329)
- ci(image-scan): ignore zlib CVE-2026-85091 until Wolfi ships the fix by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41353](https://github.com/BerriAI/litellm/pull/41353)
- feat(e2e): reuse exact provider responses for 24 hours by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41346](https://github.com/BerriAI/litellm/pull/41346)
- fix(xai): keep 'instructions' on the xAI Responses API so system messages survive web search by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;38254](https://github.com/BerriAI/litellm/pull/38254)
- feat(ui): configure capability and Fuse v2 classifiers by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41315](https://github.com/BerriAI/litellm/pull/41315)
- fix(anthropic): tolerate message\_delta events without usage when streaming by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41336](https://github.com/BerriAI/litellm/pull/41336)
- test(router): ignore deployment-selection logs in the fallback log assertion by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41358](https://github.com/BerriAI/litellm/pull/41358)
- test(proxy): assert budget resets decrement the cleared spend by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41359](https://github.com/BerriAI/litellm/pull/41359)
- fix(e2e): expect models filters to persist after reload by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41348](https://github.com/BerriAI/litellm/pull/41348)
- fix(e2e): record cookie-setting provider responses and keep prompt-caching tests live by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41366](https://github.com/BerriAI/litellm/pull/41366)
- fix(responses): recount tokens when a streamed response completes without usage by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41337](https://github.com/BerriAI/litellm/pull/41337)
- fix(ui): simplify Capability and Fuse advanced routing options by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41371](https://github.com/BerriAI/litellm/pull/41371)
- fix(mcp): authorize JWT OAuth credential persistence by [@&#8203;joshua-berri](https://github.com/joshua-berri) in [#&#8203;41314](https://github.com/BerriAI/litellm/pull/41314)
- feat(router): stream shadow traffic and fan out silent\_model to multiple targets by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41368](https://github.com/BerriAI/litellm/pull/41368)
- perf(content\_filter): scan a bounded window per streamed chunk by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41407](https://github.com/BerriAI/litellm/pull/41407)
- fix(proxy): hide model allowlist from client-facing model access denied errors by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41310](https://github.com/BerriAI/litellm/pull/41310)
- feat(http): opt-in outbound HTTP/2 for httpx clients by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41268](https://github.com/BerriAI/litellm/pull/41268)
- refactor(rust): remove gateway, config, router, realtime, and Rust trace-parity instrumentation by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41432](https://github.com/BerriAI/litellm/pull/41432)
- fix(guardrails): don't add post\_call output scan for MCP-only Presidio modes by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40571](https://github.com/BerriAI/litellm/pull/40571)
- chore(prices): sync Azure, Azure AI, Gemini, OpenAI, Bedrock, Together AI, Fireworks and Vertex prices: 278 models, 59 new, 30 deprecated by [@&#8203;berriai-litellm-provider-info-sync](https://github.com/berriai-litellm-provider-info-sync)\[bot] in [#&#8203;41154](https://github.com/BerriAI/litellm/pull/41154)
- fix(rag): forward retrieval\_filter from retrieval\_config to vector store search by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34427](https://github.com/BerriAI/litellm/pull/34427)
- refactor(rust): extract auth and cache crates by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41464](https://github.com/BerriAI/litellm/pull/41464)
- fix(proxy): default litellm\_trace\_id to the OTel server span trace id by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41386](https://github.com/BerriAI/litellm/pull/41386)
- chore(codeowners): add ryan and kerry as owners of the cost map by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41333](https://github.com/BerriAI/litellm/pull/41333)
- fix(responses): guard empty-choices chunks in the Responses API streaming bridge by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34455](https://github.com/BerriAI/litellm/pull/34455)
- chore(prices): sync Google Gemini prices: 22 models by [@&#8203;berriai-litellm-provider-info-sync](https://github.com/berriai-litellm-provider-info-sync)\[bot] in [#&#8203;41457](https://github.com/BerriAI/litellm/pull/41457)
- fix(bedrock): forward userContext in Knowledge Base Retrieve requests by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41475](https://github.com/BerriAI/litellm/pull/41475)
- ci(rust): split rust jobs, use nextest and Swatinem/rust-cache by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41480](https://github.com/BerriAI/litellm/pull/41480)
- fix(fireworks\_ai): flatten dict-form reasoning\_effort to its effort string by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41335](https://github.com/BerriAI/litellm/pull/41335)
- fix(proxy): never forward the LiteLLM virtual key to Anthropic on the /anthropic passthrough by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41340](https://github.com/BerriAI/litellm/pull/41340)
- fix(proxy): rename AWS Secrets Manager secret when key alias changes by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41468](https://github.com/BerriAI/litellm/pull/41468)
- feat(otel): promote nested request metadata keys to litellm.metadata.\* span attributes by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41462](https://github.com/BerriAI/litellm/pull/41462)
- fix(proxy): sync AWS Secrets Manager on body-less key regenerate and key alias changes by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41458](https://github.com/BerriAI/litellm/pull/41458)
- fix(http\_handler): keep a handler alive while a response it issued is still reading by [@&#8203;max-sixty](https://github.com/max-sixty) in [#&#8203;34829](https://github.com/BerriAI/litellm/pull/34829)
- fix(bedrock): make prompt caching work on the Nova InvokeModel route by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41343](https://github.com/BerriAI/litellm/pull/41343)
- ci(migrations): flag defaulted ADD COLUMN on request-log tables by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41460](https://github.com/BerriAI/litellm/pull/41460)
- feat(prometheus): add customer (end\_user) budget gauges by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41472](https://github.com/BerriAI/litellm/pull/41472)
- fix(otel): drop None metric and event attributes before OTLP export by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36815](https://github.com/BerriAI/litellm/pull/36815)
- fix(anthropic): carry the served model from message\_start onto stream chunks by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41446](https://github.com/BerriAI/litellm/pull/41446)
- fix(models): rolling registry audit: Gemini latest aliases, Nova cache pricing, OpenRouter/Together sync, Mistral GLM 5.3, Azure snapshots, Grok caching by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41112](https://github.com/BerriAI/litellm/pull/41112)
- fix(router): count TPM/RPM usage before building rate-limit headers by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41474](https://github.com/BerriAI/litellm/pull/41474)
- feat(guardrails): release buffered stream chunks after each passing scan by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41425](https://github.com/BerriAI/litellm/pull/41425)
- fix!: re-check budget on router fallback targets by [@&#8203;runjivu](https://github.com/runjivu) in [#&#8203;41379](https://github.com/BerriAI/litellm/pull/41379)
- refactor(ocr): move file preparation from the python bridge into litellm-core by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41489](https://github.com/BerriAI/litellm/pull/41489)
- feat(s3): add s3\_log\_prompts\_only option to log prompts without responses by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41327](https://github.com/BerriAI/litellm/pull/41327)
- feat(team): team-level model\_max\_budget with key-level overrides by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41330](https://github.com/BerriAI/litellm/pull/41330)
- feat(keys): filter /key/list by active, expired, revoked or deleted status and serve deleted keys from /key/info by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41311](https://github.com/BerriAI/litellm/pull/41311)
- feat(proxy): expose lifetime total\_spend on virtual keys by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41403](https://github.com/BerriAI/litellm/pull/41403)
- fix(proxy): release completed max-parallel slots promptly by [@&#8203;elifozdamar](https://github.com/elifozdamar) in [#&#8203;40843](https://github.com/BerriAI/litellm/pull/40843)
- feat(ui): accept ssh clone urls when registering a skill by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35418](https://github.com/BerriAI/litellm/pull/35418)
- fix(prices): dedupe Nova cache\_read\_input\_token\_cost keys left by a text merge by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41496](https://github.com/BerriAI/litellm/pull/41496)
- fix(otel): propagate W3C trace context on HTTP and WebSocket passthrough by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40669](https://github.com/BerriAI/litellm/pull/40669)
- fix(proxy): remove duplicate user budget hook that 429'd zero-cost models by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41345](https://github.com/BerriAI/litellm/pull/41345)
- test(logging): pick this test's own records out of the shared log batch by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41487](https://github.com/BerriAI/litellm/pull/41487)
- test(together\_ai): move request-shape checks to the mapped file, drop the live ones by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41360](https://github.com/BerriAI/litellm/pull/41360)
- feat(ui): shared URL-state layer for tables and tabs by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;41331](https://github.com/BerriAI/litellm/pull/41331)
- feat(e2e): make the provider cache reusable across builds and mount Bedrock behind it by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41402](https://github.com/BerriAI/litellm/pull/41402)
- fix(mcp): fail closed on missing upstream credentials by [@&#8203;joshua-berri](https://github.com/joshua-berri) in [#&#8203;41364](https://github.com/BerriAI/litellm/pull/41364)
- feat(rust): scaffold Redis cache crate by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41501](https://github.com/BerriAI/litellm/pull/41501)
- fix(dashscope): forward reasoning\_effort to the provider by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37506](https://github.com/BerriAI/litellm/pull/37506)
- fix(proxy): carry litellm\_call\_id through endpoint specific error logs and failure responses by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41356](https://github.com/BerriAI/litellm/pull/41356)
- fix(proxy): retry rate-limit fallbacks from a pristine request snapshot by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40596](https://github.com/BerriAI/litellm/pull/40596)
- fix(gemini): map minimal thinking to low for Gemini 3.7 and 3.8 Flash by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41201](https://github.com/BerriAI/litellm/pull/41201)
- fix(proxy): stop forwarding LiteLLM credential headers on Bedrock agent-runtime passthrough by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41504](https://github.com/BerriAI/litellm/pull/41504)
- fix(streaming): estimate interrupted Anthropic stream usage from reasoning\_content by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41503](https://github.com/BerriAI/litellm/pull/41503)
- fix(azure\_ai): route Responses API to native /openai/v1/responses for Foundry Models by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33856](https://github.com/BerriAI/litellm/pull/33856)
- fix(proxy): show all model groups to proxy admins in /model\_group/info by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41094](https://github.com/BerriAI/litellm/pull/41094)
- feat(proxy): let proxy admins choose which team fields team admins may edit by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;39996](https://github.com/BerriAI/litellm/pull/39996)
- fix(bedrock): neutralize orphaned tool blocks instead of raising or injecting a dummy tool (internal copy of [#&#8203;31400](https://github.com/BerriAI/litellm/issues/31400)) by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41513](https://github.com/BerriAI/litellm/pull/41513)
- feat(ui): persist organizations and projects list, detail tab and key table state in the URL by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;41445](https://github.com/BerriAI/litellm/pull/41445)
- fix(bedrock\_mantle): accept and forward verbosity on gpt-5.x chat completions by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41509](https://github.com/BerriAI/litellm/pull/41509)
- test: cover database transactions and persisted accounting by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41073](https://github.com/BerriAI/litellm/pull/41073)
- test: provider wire contracts, streaming and recovery by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41075](https://github.com/BerriAI/litellm/pull/41075)
- fix(mcp): count admin static headers as api\_key credential slots by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41514](https://github.com/BerriAI/litellm/pull/41514)
- ci: auto-merge provider-info-sync PRs when CI, Greptile and Bugbot are clean by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41494](https://github.com/BerriAI/litellm/pull/41494)
- feat(rust): add standalone framing crate by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41500](https://github.com/BerriAI/litellm/pull/41500)
- fix(proxy): enforce tag budgets for tags added by guardrails by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40842](https://github.com/BerriAI/litellm/pull/40842)
- fix(utils): run post-call deployment hook on converted chat streams by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41495](https://github.com/BerriAI/litellm/pull/41495)
- fix(e2e): bind provider-cache recordings to the deployment's test, not the serving process by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41520](https://github.com/BerriAI/litellm/pull/41520)
- test: add extension and browser integration contracts by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41078](https://github.com/BerriAI/litellm/pull/41078)
- fix(logging): scan each log record once and collapse base64 payloads before the secret regex by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40934](https://github.com/BerriAI/litellm/pull/40934)
- fix(spend\_tracking): attribute router-rejected requests to the model group provider by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41507](https://github.com/BerriAI/litellm/pull/41507)
- feat(router): discover token limits for hosted OpenAI-compatible models by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41508](https://github.com/BerriAI/litellm/pull/41508)
- feat(proxy): let team admins edit rpm\_limit and max\_budget when enabled by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;41525](https://github.com/BerriAI/litellm/pull/41525)
- fix(otel): fit per-index OpenInference messages to the span's remaining attribute budget by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41498](https://github.com/BerriAI/litellm/pull/41498)
- test(aws): verify rotated secret value by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41524](https://github.com/BerriAI/litellm/pull/41524)
- test(e2e): read a deleted key back as deleted, not as a 404 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41551](https://github.com/BerriAI/litellm/pull/41551)
- fix(otel v2): map the caller's Langfuse user, session and tags onto the root and generation spans by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41140](https://github.com/BerriAI/litellm/pull/41140)
- test: fix seven tests left stale by [#&#8203;41311](https://github.com/BerriAI/litellm/issues/41311), [#&#8203;41337](https://github.com/BerriAI/litellm/issues/41337), [#&#8203;39996](https://github.com/BerriAI/litellm/issues/39996), [#&#8203;41310](https://github.com/BerriAI/litellm/issues/41310), [#&#8203;41289](https://github.com/BerriAI/litellm/issues/41289) and [#&#8203;41315](https://github.com/BerriAI/litellm/issues/41315) by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41527](https://github.com/BerriAI/litellm/pull/41527)
- test(budgets): cover management null handling by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41563](https://github.com/BerriAI/litellm/pull/41563)
- test(e2e): drop the auto-router select "opens below" spec by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41568](https://github.com/BerriAI/litellm/pull/41568)
- fix(guardrails): stream Prompt Security post\_call redactions in incremental\_diff mode by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41558](https://github.com/BerriAI/litellm/pull/41558)
- fix(guardrails): give post-call scans the scoped request conversation and tools by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41220](https://github.com/BerriAI/litellm/pull/41220)
- feat(openrouter): add stealth/union-alpha to the model cost map by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41576](https://github.com/BerriAI/litellm/pull/41576)
- test(management): cover project authorization lifecycle by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41573](https://github.com/BerriAI/litellm/pull/41573)
- feat(rust): map Anthropic Messages transformations by [@&#8203;yujonglee-berri](https://github.com/yujonglee-berri) in [#&#8203;41531](https://github.com/BerriAI/litellm/pull/41531)
- fix(e2e): clear the three standing errors in the scheduled Buildkite suite by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41616](https://github.com/BerriAI/litellm/pull/41616)
- refactor(rust\_bridge): declarative route catalog and shared runtime selection by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41479](https://github.com/BerriAI/litellm/pull/41479)
- fix(mock\_completion): keep the resolved provider so router custom pricing resolves for azure\_ai deployments by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41623](https://github.com/BerriAI/litellm/pull/41623)
- fix(mcp): restrict health discovery to virtual key grants by [@&#8203;joshua-berri](https://github.com/joshua-berri) in [#&#8203;41609](https://github.com/BerriAI/litellm/pull/41609)
- fix(mcp): preserve request-selected guardrails during tool execution by [@&#8203;joshua-berri](https://github.com/joshua-berri) in [#&#8203;41619](https://github.com/BerriAI/litellm/pull/41619)
- refactor(ocr): mirror Python provider layout and preserve tests by [@&#8203;yujonglee-berri](https://github.com/yujonglee-berri) in [#&#8203;41550](https://github.com/BerriAI/litellm/pull/41550)
- test(fireworks\_ai): stop pinning vision support on minimax-m3 by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41627](https://github.com/BerriAI/litellm/pull/41627)
- perf(spend\_tracking): index LiteLLM\_SpendLogs by (api\_key, startTime) by [@&#8203;etiennechabert](https://github.com/etiennechabert) in [#&#8203;37983](https://github.com/BerriAI/litellm/pull/37983)
- fix(proxy): reject non-string model with 400 and log its spend as unknown-model by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41633](https://github.com/BerriAI/litellm/pull/41633)
- test(together\_ai): stop pinning successor deprecation status by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41635](https://github.com/BerriAI/litellm/pull/41635)
- chore(prices): sync Together AI prices: 6 models, 6 deprecated \[sync failed: Google Gemini] by [@&#8203;berriai-litellm-provider-info-sync](https://github.com/berriai-litellm-provider-info-sync)\[bot] in [#&#8203;41570](https://github.com/BerriAI/litellm/pull/41570)
- fix(budgets): page end-user cache invalidation after a budget reset by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;41488](https://github.com/BerriAI/litellm/pull/41488)
- chore: bump litellm-proxy-extras 0.4.98 -> 0.4.99 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41659](https://github.com/BerriAI/litellm/pull/41659)
- fix(tests): resolve the integration support package without run.py's PYTHONPATH by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41373](https://github.com/BerriAI/litellm/pull/41373)
- fix(ui): keep untimed guardrail entries on the request lifecycle by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41374](https://github.com/BerriAI/litellm/pull/41374)
- fix(anthropic-bridge): convert mid-conversation system turns to user turns on /v1/messages to chat completions by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41493](https://github.com/BerriAI/litellm/pull/41493)
- fix(bedrock): support aws-sdk-bedrock-runtime 0.10/0.11 in Bedrock Realtime by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41542](https://github.com/BerriAI/litellm/pull/41542)
- feat(cli): deprecate the litellm-proxy entrypoint in favour of lite by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41673](https://github.com/BerriAI/litellm/pull/41673)
- fix(scim): align pagination `count` validation with RFC 7644 by [@&#8203;zachbernstein-sdx](https://github.com/zachbernstein-sdx) in [#&#8203;41444](https://github.com/BerriAI/litellm/pull/41444)
- fix(bedrock): never emit Converse cachePoint for OpenAI-family models by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41419](https://github.com/BerriAI/litellm/pull/41419)
- fix(images): stop forwarding the raw image\[] and mask\[] form keys by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;39512](https://github.com/BerriAI/litellm/pull/39512)
- feat(management\_v1): bulk update team member budgets by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;41632](https://github.com/BerriAI/litellm/pull/41632)
- refactor(rust): extract provider translations by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41690](https://github.com/BerriAI/litellm/pull/41690)
- feat(cli): rename lite autoroute up/down to start/stop, keeping the old names as deprecated aliases by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41672](https://github.com/BerriAI/litellm/pull/41672)
- fix(responses): keep the addressed response id off bridged provider requests by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41689](https://github.com/BerriAI/litellm/pull/41689)
- fix(license): let a wildcard allowed\_features license grant the auto\_router feature by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41684](https://github.com/BerriAI/litellm/pull/41684)
- fix(ui): list every provider in the cache leakage by-model table by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40875](https://github.com/BerriAI/litellm/pull/40875)
- fix(team): keep a forked member budget's reset window and audit bulk member budget writes by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;41686](https://github.com/BerriAI/litellm/pull/41686)
- feat(proxy): add TypeSafe AI Jev evaluate passthrough with registry-priced spend tracking by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41607](https://github.com/BerriAI/litellm/pull/41607)
- test(e2e): cover bedrock batch file upload and create in the us-gov-west-1 partition by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41536](https://github.com/BerriAI/litellm/pull/41536)
- feat(grafana): add all-metrics dashboard and fix stale dashboard\_v2 gauges by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41578](https://github.com/BerriAI/litellm/pull/41578)
- fix(cost): price Azure PTU spillover requests at standard token rates by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41569](https://github.com/BerriAI/litellm/pull/41569)
- build(deps): bump soupsieve to 2.9.2 to clear the osv-scan advisories by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41703](https://github.com/BerriAI/litellm/pull/41703)
- fix(fireworks\_ai): restore supports\_vision on minimax-m3 in the cost map by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41699](https://github.com/BerriAI/litellm/pull/41699)
- feat(policy\_engine): explicit priority for policy attachment execution order by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41571](https://github.com/BerriAI/litellm/pull/41571)
- feat(router): add TypeSafe Jev as a complexity router classifier by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41615](https://github.com…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants