feat(policy_engine): explicit priority for policy attachment execution order - #41571
Conversation
…n order Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
🤖 Devin AI EngineerI'll be helping with this pull request! Here's what you should know: ✅ I will automatically:
Note: I can only respond to comments from users who have write access to this repository. ⚙️ Control Options:
|
|
|
|
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
…in UI Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
Admin UI proof at 669a664: Attachments tab before and after, priority form validation, and the Policy Simulator showing the new order Before (5ef40a6), Attachments tab has no Priority column: After (669a664), Attachments tab with the Priority column: After, out of range value rejected in the form: After, non integer value rejected in the form: After, Policy Simulator for gpt-4o-mini with tag production, model-policy listed before tag-policy: |
|
bugbot run |
… by keystroke Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 8164189. Configure here.
|
Audit UI evidence at 8164189 (before at 5ef40a6): attachments table, add form validation, simulator order, Logs page, and the annotated recording
|
TLDR
Problem this solves:
How it solves it:
priorityon a policy attachment (config, API, DB, Admin UI)prioritykeep today's tier orderpriorityis bounded to a signed 32-bit integer, so bad values fail with 422User Flow
Before: an admin wants the model-scoped policy to run first but the tag-scoped policy always wins
tag-policywithtags: [production]andmodel-policywithmodels: [gpt-4o-mini]in config.yaml{"model": "gpt-4o-mini", "tags": ["production"]}matched_policiescomes back astag-policythenmodel-policy, and there is no field to change thatx-litellm-applied-policies: tag-policy,model-policyAfter: the admin sets
priorityand the order follows itpriority: 1to themodelsattachment andpriority: 2to thetagsattachmentmatched_policiescomes back asmodel-policythentag-policyx-litellm-applied-policies: model-policy,tag-policyRelevant issues
Affected release
Linear ticket
Resolves LIT-7979
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Live audit (no mocks) of the merge base 5ef40a6 and the PR tip 8164189, each in its own worktree with its own Postgres database, two uvicorn workers,
LITELLM_DISABLE_NO_REDIS_WARNING=true, launched withpython litellm/proxy/proxy_cli.py --config audit_config.yaml --port <4101 base | 4100 head> --num_workers 2 --use_v2_migration_resolver. Real OpenAI (openai/gpt-5.6) and Anthropic (anthropic/claude-sonnet-5) calls. Every LLM cell was also checked inLiteLLM_SpendLogsbyrequest_id = <response id>(/v1/responsesstreams fall back to thex-litellm-call-idstored in metadata, pre-existing). The same config ran on both legs; the base leg silently ignorespriorityin it. Screenshots and the annotated recording are in this comment, the full matrix (SDK sync and async cells, streaming, sad paths, edge cases, chaos) is attached to the Devin session asaudit_report.mddb-policy(no guardrails) was created on both legs withPOST /policiesso DB attachments could be added through the API and the Admin UIBefore (5ef40a6)
Resolve order
curl -s -X POST http://localhost:4101/policies/resolve -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model":"audit-ordered","tags":["ordered"]}'{"matched_policies":[{"policy_name":"block-tag-policy","matched_via":"tag:ordered"},{"policy_name":"model-policy","matched_via":"model:audit-ordered"},{"policy_name":"block-model-policy","matched_via":"model:audit-ordered"}]}Tag tier runs before model tier no matter what the config asked for
Blocking pipeline order on a real request
curl -s -D - http://localhost:4101/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -H 'x-litellm-tags: ordered' -d '{"model":"audit-ordered","messages":[{"role":"user","content":"Reply with the single word ok"}],"max_tokens":5}'The tag-scoped pipeline blocks first, so the model-scoped one never gets to run
/v1/chat/completions
curl -s -D - http://localhost:4101/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -H 'x-litellm-tags: production' -d '{"model":"audit-openai","messages":[{"role":"user","content":"Reply with the single word ok"}],"max_completion_tokens":16}'psql "$DATABASE_URL" -c "select request_id from \"LiteLLM_SpendLogs\" where request_id='chatcmpl-EPDgUqMmWIloGzl3zOa6eVxHobKpB'"returns 1 row/v1/messages
curl -s -D - http://localhost:4101/v1/messages -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -H 'x-litellm-tags: production' -d '{"model":"audit-anthropic","max_tokens":16,"messages":[{"role":"user","content":"Reply with the single word ok"}]}'LiteLLM_SpendLogshas 1 row formsg_011Cf9hYXdfuyiGYG8mHPnVK/v1/responses
curl -s -D - http://localhost:4101/v1/responses -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -H 'x-litellm-tags: production' -d '{"model":"audit-openai","input":"Reply with the single word ok","max_output_tokens":16}'The merge base returns a provider 404 for this model on the Responses API (reproduced twice, unrelated to policies; the policy headers still show the old tag-first order). The same request returns 200 on the PR tip below
Attachment create and list
curl -s -X POST http://localhost:4101/policies/attachments -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"policy_name":"db-policy","tags":["api-created"],"priority":5}'{"attachment_id":"ccbfbcca-e303-467b-a9d6-8f67c7e67068","policy_name":"db-policy","tags":["api-created"],"created_at":"2026-09-17T21:15:46.142000Z","definition_location":"db"}priorityis silently dropped, the response andGET /policies/attachments/listhave no such fieldOut of range priority
priorityfield to validateAdmin UI
Chaos
Spend tracking - transient DB error writing spend logs, retry 3/3Guardrail modes, callback modes, Claude Code
Same commands as the After section below, run against port 4101. Per-request, key attached and team attached guardrails were all added on top of the policy guardrails, the
generic_apicallback fired for YAML, key level and team level success and failure exactly as on head, and the Claude Code TUI answeredpriority audit. Onlyx-litellm-applied-policiesdiffers: the base runs the fixed tier order (tag-policy,model-policyand, with the team key,team-policy,tag-policy,model-policy)After (8164189)
Resolve order
curl -s -X POST http://localhost:4100/policies/resolve -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model":"audit-ordered","tags":["ordered"]}'{"matched_policies":[{"policy_name":"model-policy","matched_via":"model:audit-ordered"},{"policy_name":"block-model-policy","matched_via":"model:audit-ordered"},{"policy_name":"block-tag-policy","matched_via":"tag:ordered"}]}Priority 1, then 10, then 20: the model-scoped pipeline now runs before the tag-scoped one
Blocking pipeline order on a real request
Same curl against port 4100
The model-scoped pipeline blocks first, which is what LIT-7979 asks for. Same result on
/v1/messagesand/v1/responseswith the same body shape/v1/chat/completions
Same curl against port 4100
LiteLLM_SpendLogshas 1 row forchatcmpl-EPDgKg5RY4vYP1zRLXaNGqMB7Ch8K. Streaming curl (chatcmpl-EPDgLaZtWx6rukY5DBnOm79CvTBCL) and the openai SDK sync, sync stream, async and async stream cells all returned the same headers and one spend log row each/v1/messages
Same curl against port 4100
1 spend log row for
msg_011Cf9hXuZVyLSGcABNTyuoT; streaming curl and the anthropic SDK sync, sync stream, async and async stream cells match/v1/responses
Same curl against port 4100
1 spend log row for that id; streaming curl and the openai SDK sync, sync stream, async and async stream cells match
Attachment create and list
curl -s -X POST http://localhost:4100/policies/attachments -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"policy_name":"db-policy","tags":["api-created"],"priority":5}'{"attachment_id":"291fc75c-8c5f-4a53-b534-3083b822842f","policy_name":"db-policy","tags":["api-created"],"priority":5,"definition_location":"db"}curl -s http://localhost:4100/policies/attachments/list -H 'Authorization: Bearer sk-1234'listspriorityon every attachment: config2, 1, null, 10, 20and the DB row5curl -s -X POST http://localhost:4100/policies/resolve ... -d '{"model":"audit-openai","tags":["production","api-created"],"team_alias":"platform"}'returnsmodel-policy, tag-policy, db-policy, team-policy: priorities 1, 2, 5, then the unprioritised team attachment lastOmitted
priorityand explicit"priority": nullboth store SQL NULL and sort as unprioritised;0,-5,-2147483648and2147483647round-trip exactly through the DB and sort numerically (psqlrows in the audit report)Out of range priority
curl -s -w '\nHTTP %{http_code}' -X POST http://localhost:4100/policies/attachments -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"policy_name":"db-policy","tags":["sad"],"priority":2147483648}'-2147483649gives 422greater_than_equal,"high"gives 422int_parsing,1.5gives 422int_from_float, no row is written.priority: highorpriority: 2.5in config.yaml stops the proxy at boot with the pydantic message instead of being silently ignoredAdmin UI
db-policy, tagui-created, Priority3, Create Attachment: the row appears with Priority 3 andGET /policies/attachments/listreturns"priority": 39999999999in Priority: the form showsPriority must be at most 2147483647and does not submitaudit-ordered, tagordered, Simulate:model-policy,block-model-policy,block-tag-policyin that orderchatcmpl-EPDgKg5RY4vYP1zRLXaNGqMB7Ch8Kand the audited/v1/messagesand/v1/responsesrows as successful requestsChaos
db: disconnectedwhile down anddb: connectedafter restart. 28 of 30 spend logs landed exactly once, 2 were dropped afterretry 3/3, the same pre-existing loss as the base leg above. The PR does not touch the spend writer, so this is follow-up material, not a regressionkill -9mid burst: the other worker kept serving, a new worker came up within 3 s, in-flight requests on the dead worker were lost as expectedGuardrail modes, callback modes, Claude Code (outside the diff, expected unchanged from base)
/key/generatewithmetadata.guardrails: [team-guardrail]) on/v1/messages, and team attached guardrail (/team/newaliasplatformwithmetadata.guardrails: [tag-guardrail], key in that team) on/v1/responsesand on streaming/v1/chat/completions: all 200, the attached guardrail appears inx-litellm-applied-guardrails, and the team key addsteam-policy=team:platformtox-litellm-policy-sourcesbehind the two prioritised policies (x-litellm-applied-policies: model-policy,tag-policy,team-policy)generic_apiHTTP sink (GENERIC_LOGGER_ENDPOINT): YAMLsuccess_callbackandfailure_callbackon a second head proxy, key levelmetadata.logging: [{callback_name: generic_api, callback_type: success}], team levelPOST /team/{team_id}/callbackwithcallback_type: failure, and a key with no callbacks. Success events arrived only where a success callback was configured, failure events (unknown model, 400 returned to the caller) only where a failure callback was, the uncallbacked key produced none, andPOST /key/healthreported{"callbacks":["generic_api"],"status":"healthy"}. Every request has exactly oneLiteLLM_SpendLogsrow by call idANTHROPIC_BASE_URL=http://localhost:4100,ANTHROPIC_MODEL=audit-anthropicand a virtual key taggedproduction: the TUI answeredpriority audit, and both requests it made landed inLiteLLM_SpendLogswithapplied_guardrailsfrom the tag and model policiesType
🆕 New Feature
Caveats (if any)
Medium
priorityonLiteLLM_PolicyAttachmentTable, additive migration, no row rewritesLow
matched_viaeffective_guardrailson/policies/resolvestays alphabetically sorted; guardrails in the flat list run in the proxy's callback order (the resolver unions them into a set, unchanged from main), sopriorityorders pipelines, thematched_policieslist and thex-litellm-applied-policiesheader, not the flat list orx-litellm-applied-guardrailsPUT/PATCH /policies/attachments/{id}on main, so a priority is changed by deleting and recreating the attachment; SpendLogs metadata carriesapplied_guardrailsonly, so policy names and order are visible in the response headers and the simulator, not the Logs page. Both are follow-ups outside this ticketFinal Attestation
ran /live-pr-risk and found no regressions/backward incompatible risks
Link to Devin session: https://app.devin.ai/sessions/2357786e068841ec904e342e4e160b2f
Open in Devin Desktop: https://app.devin.ai/desktop/session/2357786e068841ec904e342e4e160b2f?variant=devin
Requested by: @yucheng-berri