fix(transcription): stop a zero output rate from zeroing transcription cost - #36914
Merged
yassin-berriai merged 1 commit intoAug 14, 2026
Conversation
Contributor
Greptile SummaryThe PR makes per-second transcription pricing fall back to a billable input rate when the configured output rate is zero.
Confidence Score: 5/5The PR appears safe to merge. No blocking failure remains.
|
| Filename | Overview |
|---|---|
| litellm/llms/openai/cost_calculation.py | Updates per-second cost selection so a zero output rate no longer suppresses a positive input rate. |
| tests/test_litellm/llms/openai/test_cost_calculation.py | Adds focused regression tests for per-second transcription pricing and representative built-in models. |
Reviews (2): Last reviewed commit: "fix(transcription): stop a zero output r..." | Re-trigger Greptile
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
…n cost cost_per_second treated a declared-but-zero output_cost_per_second as a real rate, so the output branch claimed the call and the elif locked out input_cost_per_second. Every transcription model shipping output_cost_per_second 0.0 next to a real input rate billed $0, which covers 43 of the 55 per-second entries in the cost map: all 36 deepgram models, both assemblyai, both elevenlabs scribe, both groq whisper and azure-stt. Custom deployments pairing the two fields the same way billed $0 as well Take the output branch only when that rate is actually billable, so a zero falls through to the input rate. Entries that duplicate one rate into both fields, whisper-1 among them, keep billing exactly what they bill today
hMED22
force-pushed
the
litellm_fix_transcription_zero_output_rate
branch
from
August 14, 2026 19:34
436a350 to
24406d6
Compare
Contributor
Author
|
@greptileai please re-review the current head 24406d6. It rebases onto latest staging so the flaked core-utils shard reruns clean |
yassin-berriai
approved these changes
Aug 14, 2026
yassin-berriai
merged commit Aug 14, 2026
29fe342
into
BerriAI:litellm_internal_staging
72 checks passed
AriOliv
added a commit
to AriOliv/litellm
that referenced
this pull request
Aug 31, 2026
* feat(lint): exempt TypedDict-annotated dict literals from LIT002 * fix(scripts): unwrap PEP 604 unions in LIT002 TypedDict detection * fix(mcp): expose client HTTP headers to logging callbacks and hooks (#36724) * fix(mcp): expose client HTTP headers to logging callbacks and hooks MCP protocol tool calls built a synthetic Request with only content-type, so metadata.headers reaching logging callbacks and guardrails was empty while /mcp-rest/tools/call exposed the full set. Rebuild the synthetic request from the connection's raw headers (shared with the sampling path), and pass sanitized headers to the pre-call hook, the MCP to LLM guardrail bridge and the Responses API MCP bridge. Credential headers stay masked and proxy key headers stripped. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): strip custom proxy key and upstream MCP credential headers from logging copies Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(mcp): make client side auth header name accessor public Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): strip custom proxy key and client redaction opt-out from mcp headers Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): drop custom proxy key header in the synthetic request builder Strips general_settings.litellm_key_header_name in build_synthetic_mcp_request so every caller, including sampling, is covered, and reverts passing general_settings into add_litellm_data_to_request on the tool call path since that also switches on enforced_params. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: shivam <[email protected]> * test(proxy): stop monkeypatch.undo re-planting fixture-mocked prisma_client * fix(ptu): stop per-token billing on a PTU-configured deployment (#36829) A deployment with PTU flat-cost attribution also billed every request per token, so a team paid for reserved capacity and again for the traffic that capacity serves. Nothing set the per-token price and an unset price falls back to the public cost map, which made the double charge the default. /model/new and /model/{id}/update now store zero for every pricing field the cost map could otherwise fill, refuse a price the caller supplies alongside PTU config with a 400 naming the field, zero a price already on the row rather than rejecting later edits of unrelated fields, and drop the zeros again when the PTU config goes. A PTU deployment is no longer read as a free model by the budget checks, which would have waived every budget for it. * fix(batches): mark terminal batch with no output file as processed in CheckBatchCost A managed batch whose request lines all failed can reach a terminal provider status (completed) with output_file_id=None and only an error_file_id. Such a row matched neither the completed-with-output billing branch nor the failed/expired/cancelled branch, so batch_processed stayed False and the poller re-selected it on every cycle for the lifetime of the deployment; output/error file deletion is also gated on batch_processed, so those files could never be deleted. Broaden the terminal handling so a completed/complete/expired batch with an output file is billed, and any terminal batch with nothing to bill (failed/cancelled, or completed/expired with no output) is marked terminal exactly once. Non-terminal statuses (validating/in_progress) are still left for the next poll, and an expired batch that did produce output is now billed. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): track spend for OpenAI passthrough /v1/embeddings (#36660) * fix(proxy): track spend for OpenAI passthrough /v1/embeddings OpenAI passthrough embeddings returned 200 but wrote no spend because the route was unsupported and Cohere's /v1/embed prefix stole the match. * fix(proxy): clear embeddings lint and Greptile comment nits Inline embeddings cost tracking to avoid new LIT001/002 hits, trim redundant doc comments, and cover the Cohere /v1/embeddings collision. * fix(proxy): drop unreachable embeddings TypeError guard convert_to_model_response_object with response_type=embedding already returns EmbeddingResponse; the isinstance check was dead patch coverage. * fix(batches): persist real terminal status when billing expired batches * fix(access groups): sync assigned_team_ids from the team write paths (#36825) * fix(spend): give a batch's cost row a primary key of its own request_id is the primary key of LiteLLM_SpendLogs and the flush inserts with skip_duplicates, so a spend log whose id already exists is dropped with no error raised and a "processed 1 spend log" line still logged. Batch cost accounting produced exactly such an id twice over, and on a proxy with message redaction enabled no batch cost row could be written at all. get_spend_logs_id derived the id by md5-hashing the response for two call types, aretrieve_batch and acreate_file. Redaction makes that hash a constant: perform_redaction returns the fixed {"text": "redacted-by-litellm"} placeholder for any shape it cannot redact, which is what a batch object and a file body both become, so every such row hashed to md5('{"text": "redacted-by-litellm"}') = 00fcbef15a3b0097e14b0ca016ed30a0 regardless of provider, user, or amount. The first row to claim that id owned it and every later row was discarded. Verified against a live proxy: four payloads spanning two providers and three distinct spend values all computed that id, and the table held one acreate_file row dating to 2025-05-25, the row that had claimed it. Keying off the batch's own identity instead is necessary but not sufficient, because creating a batch already writes an acreate_batch row under exactly that id, so the cost row becomes a duplicate of the batch's own creation row. Also verified live: after the hash was removed the poller computed and flushed a batch's cost, and the only row carrying that id was the acreate_batch row from when the batch was submitted. The id now comes from the response's own id, then the standard logging payload's id, then litellm_call_id, and a batch cost row is namespaced with a _batch_cost suffix so it cannot collide with the creation row. The middle term is what keeps this correct under redaction: that payload is built from the unredacted response, so it still carries the batch id after redaction has flattened the body. Keying the cost row to the batch rather than to the call also keeps accounting the same batch twice collapsing to one row instead of billing it twice. Every other call type still derives its key exactly as before. Cost and usage themselves are unaffected by redaction: the token columns fall back to the standard logging payload and spend comes from its response_cost, neither of which redaction touches. generate_hash_from_response had no other caller and is removed with it. * test(spend): annotate the batch cost row constants as Final * fix(bedrock): resolve the managed-batch output bucket on the model-routed and cost-poller paths get_configured_s3_bucket_name accepts the output bucket only from the immutable _litellm_internal_model_credentials snapshot or AWS_S3_BUCKET_NAME. That refusal to read litellm_params is deliberate: the bucket is what validate_managed_cloud_file_id checks a file id against, so trusting a request-supplied value would let a caller redirect reads to a bucket of their choosing Two live entry points reach the Bedrock file-content transformation without ever building that snapshot. The managed-files pre-call hook sets data["model"] for any id carrying llm_output_file_id, which is every batch output, so get_file_content always takes the model-routed branch; that branch called llm_router.afile_content directly, and managed_files_obj.afile_content, the only caller that built the snapshot, is therefore unreachable for batch output. CheckBatchCost spread the deployment credentials as plain kwargs, and get_litellm_params does not carry s3_bucket_name across (gcs_bucket_name is listed for exactly this reason, its S3 counterpart is not), so the poller lost the bucket the same way The result was that every completed Bedrock managed batch failed files.content with "S3 bucket_name is required" and never had its cost tracked, leaving the row to be re-polled every cycle. Both paths now resolve the deployment credentials and pass the same MappingProxyType snapshot the managed-files hook already builds * test(files): capture routed retrieval calls immutably The mock merged every call into one shared dict, so a second routed retrieval would overwrite the first and the assertions would still pass. Keep one frozen snapshot per call and assert exactly one call, which also makes an unintended second retrieval a failure rather than something the merge hides * fix(bedrock): resolve the managed-batch output bucket on the inline accounting path too A third path reads a completed batch's output file, and it could not resolve the bucket either. When cost is accounted from the retrieve itself rather than from the poller, the batch success handler calls _handle_completed_batch, which fetches the output file through _extract_file_access_credentials. That helper forwarded a whitelist covering Azure and Vertex, gcs_bucket_name included, but nothing for Bedrock, and retrieve_batch built its litellm_params through get_litellm_params, whose fixed signature drops the trusted credential snapshot. So the snapshot never reached the file read and it failed with "S3 bucket_name is required" for a bucket the deployment had configured, leaving the batch's cost unrecorded. Adding s3_bucket_name to that whitelist would not have worked. The Bedrock file config deliberately resolves the bucket only from the immutable server-side snapshot or the environment, never from a request param, because the bucket is what managed file ids are validated against. The snapshot is therefore what has to flow, exactly as it already does for the model-routed and cost-poller paths. retrieve_batch now re-adds the snapshot after get_litellm_params, the same way the file operations already do, the whitelist forwards it, and the proxy attaches it for router-routed managed batches from the deployment behind the unified id. Verified against a live proxy reading a real completed Bedrock batch: the cost row appears within seconds of the retrieve carrying the batch's real spend and usage, where before the read raised and no row was written. Resolving those credentials is best effort. A batch whose deployment no longer resolves, which happens when a model group is removed while batches are in flight, still serves its status instead of failing the request on the lookup. This matters for the OSS and polling-disabled configurations, where the retrieve path is the only thing that accounts for a batch at all. * refactor(batches): share the trusted-credentials helper across both call paths The helper that carries the credential snapshot into litellm_params lived private in files/main.py, and the batch retrieve needed it too. It now sits beside get_litellm_params, which is what it augments, so neither caller reaches into the other's private surface. Typed as Mapping/MutableMapping of object rather than Any, which the strict import rules ban. The file-content route builds the snapshot through the same helper as the batch route instead of assembling a conditional mapping inline, which drops two mutable constructions and leaves one way to attach it. Its name loses the batch suffix now that both routes use it. * fix(batches): account a managed batch's cost exactly once Two components computed a managed batch's cost and each assumed it was the only one. Retrieving a batch computed it through the @client decorator's success callback, and CheckBatchCost computed it on its own schedule. Whichever observed completion first decided the outcome, so cost was either counted once per retrieve or not at all. The lockout is the worse half. Retrieving a batch that had reached completion set batch_processed=True, which is what takes a batch out of CheckBatchCost's queue, since it selects batch_processed=False. That write claimed the cost had been accounted for on behalf of a callback that had not run yet and was not awaited. When the callback then failed the cost was gone permanently, with the poller already retired and no retry left. Observed on a live proxy: two completed batches whose callbacks raised inside the logging worker, one on a provider output path that did not resolve and one on a batch whose output file id was still None, both left marked processed with no spend row and no way to recover them. Nothing logged at error level for the batches themselves. The over-count is the other half. Nothing suppressed recomputation, so each retrieve of an already-completed batch recorded that batch's full cost again. A caller polling its own batch to see whether it had finished inflated spend by however many times it looked. The flag now means what its name says, and only the component that actually recorded the cost sets it. When the poller is running it owns accounting, so retrieving a managed batch records no cost and leaves the flag alone; the poller computes once and sets it. When the poller cannot be relied on, either because polling is disabled by config or because the enterprise job never registered, the retrieve path is the only accountant and behaves exactly as before. Batches with no managed object row are untouched either way, since neither the flag nor the poller queue applies to them. * fix(batches): only hand accounting to the poller once it can mark batches done The handoff asked whether the poller was running, when what matters is whether it will actually account for the batch. Those differ on a schema without the batch_processed column: the poller cannot filter on it, so it falls back to a query that excludes complete and completed rows, and it cannot set it either. A caller retrieving a provider-completed batch before the poller saw it therefore suppressed inline accounting, then marked the row complete, and the fallback query could never find it again. Nobody accounted for that batch, so its cost escaped the caller's budget entirely. The poller now publishes batch_processed_support_confirmed, set only once a filtered query has actually succeeded, and the handoff requires it. Defaulting to unconfirmed keeps accounting on the retrieve path in exactly the cases the poller would drop the batch, including the window before the poller's first cycle. All four combinations account exactly once: unconfirmed leaves the retrieve accounting and setting the marker, whether or not the column exists, and confirmed is only reachable when the column is present, where the poller accounts and sets it. A scheduler that hands back something other than a bound method leaves no poller to interrogate, which reads as unconfirmed rather than as working. * fix(batches): decide batch cost ownership once per retrieve The ownership question was asked twice for one retrieve: once before the provider call to decide whether to suppress inline accounting, and again afterwards to decide whether to mark the batch accounted. Between those two points the poller can complete its first successful filtered query and become usable, so the two answers disagree. The retrieve then accounts for the batch inline, having decided the poller was unusable, while the later check sees a usable poller and leaves the marker unset, so the poller accounts for the same batch again and its spend is counted twice. The retrieve now decides once and passes that decision to update_batch_in_database, which prefers it over re-deriving one. Callers that record no cost of their own leave it unset and keep deriving it as before, so the cancel path is unchanged. * ci: drop the CircleCI ui_build and ui_unit_tests jobs (#36893) Both are covered on GitHub Actions. test-litellm-ui-build.yml runs the dashboard build on every PR, and test-litellm-ui-unit.yml runs the vitest suite with ui-unit-tests already a required check, so neither CircleCI job gates anything that GHA does not already gate. ui_build additionally produced nothing anyone consumed. It persisted litellm/proxy/_experimental/out to the workspace, and the only job downstream of it was ui_unit_tests, which never attached the workspace and reinstalled from source instead. The requires edge was pure sequencing, so the build output was written and discarded on every client-touching PR. One real narrowing comes with this, and it is deliberate. ui_unit_tests ran the full vitest suite on PRs, while the GHA job scopes PR runs to tests reachable from the diff and keeps the full suite on pushes to staging. That split was a measured decision in #34175 and it still holds: the suite is 252s and 248s of that is CreateMCPServer.integration.test.tsx alone, so running everything per PR buys about four minutes to re-run one file. Note that assert-ci-coverage does not speak to this. It walks tests/**/test_*.py only, so it is blind to vitest files by construction; it stays green here because no Python test lost a runner, which is a narrower claim than the UI side being unaffected. auth_ui_unit_tests is a different job, a Python suite on a Postgres sidecar, and is untouched * fix(langfuse): source the emitted metadata blob from StandardLoggingPayload (#36744) Request metadata carries the whole UserAPIKeyAuth object, whose team_metadata holds the customer's own langfuse callback_vars. The only filter on the emitted blob was a four key deny list written as a circular reference crash guard, so those credentials reached the customer's own langfuse traces. The emitted blob is now the StandardLoggingPayload allowlist plus the litellm computed enrichments, and nothing is copied across from raw request metadata. That makes the credential exclusion structural rather than a filter someone has to keep correct. Steering keys keep reading raw metadata, matching literal_ai. Proxy callers are unaffected: their request metadata already rides under the allowlisted requester_metadata key, nesting intact. debug_langfuse dumped raw request metadata into the trace as a second copy of the same leak. It now emits caller scalars only. When StandardLoggingPayload is absent the trace is still emitted with the existing trace_id fallback, so failure traces survive. * refactor(ui): migrate Navbar off antd to shadcn Replaces Ant Design with the in-repo shadcn layer across every Navbar component, removing the last antd imports from src/components/Navbar. - CommunityEngagementButtons, NotificationsBell, ViewSwitcher, BlogDropdown, WorkerDropdown and UserDropdown now compose @/components/ui primitives - antd icons render at 1em while lucide defaults to 24px, so every icon carries an explicit size class matching what it replaced - UserDropdown uses Popover rather than DropdownMenu: its panel holds switches and badges, and form controls inside role="menu" are invalid - WorkerDropdown moves to Combobox since shadcn Select has no search - drops the nine no-restricted-imports suppressions these files no longer need * refactor(ui): migrate log details drawer off antd to shadcn Replaces Ant Design across every source file under src/components/view_logs, so the request log drawer and its viewers compose @/components/ui primitives. - Drawer becomes Sheet, Collapse becomes Collapsible, Segmented and Radio.Group become Tabs, Tag becomes Badge, Descriptions becomes a local grid helper - every lucide icon carries an explicit size class, since antd icons render at 1em while lucide defaults to 24px - two tests dropped assertions on antd internal class names in favour of rendered text and roles, and the Pretty/JSON case now proves the toggle actually swaps the body rather than only that both controls render - drops the eslint suppressions these files no longer need * fix: report real token usage on guardrail-blocked /v1/responses replies ## TLDR Signed-off-by: Ishaan <[email protected]> * refactor(ui): migrate AI Hub off antd and tremor to shadcn Replaces Ant Design and Tremor across src/components/AIHub, so the model, agent, MCP and skill hub views compose @/components/ui primitives. - Modal becomes Dialog, tremor TabGroup becomes Tabs, tremor Card and Table become their shadcn counterparts, and Tag and tremor Badge become Badge - the three publish forms wrapped antd Form around zero Form.Item fields, so the wrapper became a div and the dead useForm and resetFields calls went with it, rather than pulling in react-hook-form for a form with no fields - antd Steps has no shadcn equivalent, so each form inlines a small ol stepper - cells holding model names, server ids and URLs gained min-w-0 and break-words so a long value cannot bleed into the neighbouring column - the three form tests dropped assertions invented by their antd mocks in favour of roles and rendered text - drops the eslint suppressions these files no longer need * fix(ui): announce the account popover as a dialog, not a menu The panel holds switches and ordinary buttons rather than menu items, so menu semantics promised keyboard behavior it does not provide. * fix(ui): give the request details drawer an accessible name Screen readers announced an unnamed dialog. The visible header is a custom layout, so the title is visually hidden to keep the drawer layout unchanged. * refactor(ui): migrate shared common_components off antd and tremor Replaces Ant Design and Tremor in the six shared components under src/components/common_components, which between them are reached by nine routes. - antd Table becomes the ui/table primitives, and the Actions column keeps antd's fixed: "right" behaviour via a sticky cell - Tremor Icon, Text and Badge become a plain span, p and StatusBadge - antd Tooltip and Typography copyable become the shadcn Tooltip and the shared CopyButton - every public prop signature is unchanged, since these are shared components and a renamed prop would break callers far from this folder - two tests dropped assertions on antd internal class names and on DOM structure, and gained cases proving a disabled action does not fire onClick MemberTable keeps a type-only import of antd's ColumnsType because a consumer annotates its own column array with it. No antd code ships from the file. * test(ui): assert the publish button is disabled while submitting The migration closed a double submit hole that antd left open, but the rewritten tests only proved the flow had not completed, so removing the guard would not have failed them. Verified by mutation: dropping disabled={loading} fails exactly this case. * refactor(ui): migrate key info and permissions views off antd and tremor Replaces Ant Design and Tremor in the key info header and detail view, the agent and vector store permission panels, and the team member permissions table. - antd Popover, Dropdown and Modal become HoverCard, DropdownMenu and Dialog, and Tremor TabGroup becomes Tabs with keepMounted so panel state survives a tab switch the way Tremor's did - the key id copy control moves to the shared CopyButton, which also fixes an icon that rendered at 24px because it inherited the heading font size - antd Checkbox onChange becomes onCheckedChange - every public prop signature is unchanged, since these are shared views - three member permission tests were passing vacuously: they searched for an unchecked box by reading .checked, which is undefined on a Base UI checkbox, so the assertions sat inside an if that never ran. They now scope the checkbox to its own row and assert the toggle, the save and the revert - drops the eslint suppressions these files no longer need * refactor(ui): migrate router settings and shared badges off antd and tremor Replaces Ant Design and Tremor in the fallbacks views, the router general settings panel, and the two shared banner and badge components. - Tremor Card, Table and Icon become the ui/card, ui/table and lucide equivalents, reproducing Tremor's icon box so click targets keep their size - antd Alert becomes a composed role="alert" region, since the shadcn CLI's alert pulls in class-variance-authority, which this repo does not have - antd InputNumber becomes a native number input, and Switch onChange becomes onCheckedChange - shadcn TableCell ships whitespace-nowrap where Tremor's did not, so cells holding model names and setting descriptions get whitespace-normal back - adds a DeprecationBanner test covering naming, the link, and dismissal, proven against the antd version first and mutation checked - drops the eslint suppressions these files no longer need * refactor(ui): move the model hub and model select onto shadcn primitives Rebuilds public_model_hub, MakeSkillPublicForm, ModelSelect and the guardrail LogViewer on the in-repo shadcn layer, so they inherit the dashboard's design tokens instead of styling themselves through Ant Design and Tremor. Public prop signatures are unchanged, so no caller moves. The two teams e2e steps that reached into antd's Select internals now drive the combobox through its test id, role and data-slot instead. * refactor(ui): give MemberTable its own extra-column type extraColumns was typed as antd's ColumnsType while the adapter only honoured string/ReactNode titles, plain-string dataIndex values and element/string/number render results, so several valid antd column forms produced blank cells. MemberTableColumn now describes exactly what the table renders, and a column with a dataIndex but no render falls back to the member value instead of rendering nothing. * refactor(ui): move the shared dropdowns and selectors onto shadcn primitives Rebuilds the thirteen form-free components under common_components on the in-repo shadcn layer, so they inherit the dashboard's design tokens instead of styling themselves through Ant Design and Tremor. SearchSelect and the three dropdowns that wrap it now forward an optional input id, so an antd Form.Item label still resolves to its control. The e2e steps that reached into antd's Select and Modal internals now go through the test id, role and data-slot. * fix(model_prices): correct Gemini 2.5 shutdown dates and DeepSeek V4 max output tokens Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(ui): cover appending a second model in ModelSelect The rewritten suite only ever picked one ordinary model, so a regression that replaced the selection instead of appending to it would have gone unnoticed. The case passes against the antd version too, so it pins behavior the migration preserves rather than adds. * test(ui): spread the real lucide-react module in the KeyInfoView mock The mock returned only CopyIcon and CheckIcon, so any icon a child later imports resolves to undefined. DeleteResourceModal now renders CircleAlert, which broke all twelve cases in this file. * fix(ui): hold the delete dialog open mid-deletion and keep unmatched select values DeleteResourceModal let escape, the backdrop and the close button dismiss it while the delete request was still in flight. SearchSelect blanked its field whenever the value was missing from options, which happens while they load; it now falls back to the raw value the way PaginatedSearchSelect already did. * refactor(ui): move the root-level dashboard components onto shadcn primitives Rebuilds nine components under src/components on the in-repo shadcn layer: both banners, the navbar chrome, the onboarding link dialog, the model filters, the model group alias table, the object permissions and logging settings views, and the user dashboard grid. Every public prop signature is unchanged, so no caller moves. * revert(ui): keep the onboarding link modal on antd The invitation dialog opens over the still-antd Invite User modal. Lifting only the shadcn dialog content above antd's mask leaves its own backdrop underneath, so an outside click reaches the wrong modal. Adding a second backdrop stops that but does not restore dismissal, and the same hazard already ships in three guardrails modals, so the stacking needs one shared fix rather than a fourth local workaround. * fix(databricks): surface prompt-cache token counts in streaming usage chunk_parser built ModelResponseStream without passing usage, so the cache_read_input_tokens and cache_creation_input_tokens that Databricks returns for Anthropic models never reached the cost calculator. Every streamed request was billed at the full input rate even when served from cache. ModelResponseStream already coerces a usage dict into Usage, which maps those keys into prompt_tokens_details, so passing the chunk's usage through is sufficient. * refactor(ui): move the settings page and bulk user invite onto shadcn primitives Rebuilds settings.tsx and bulk_create_users_button.tsx on the in-repo shadcn layer. The settings callback form moves from antd Form to react-hook-form with the shared Field primitives, and the CSV drop zone replaces antd Upload with a native file input plus drag handlers. Both public prop signatures are unchanged, so no caller moves. * Remove comment about prompt-cache usage in test Remove outdated comment regarding prompt-cache counts in chunk_parser. * refactor(ui): move the cost tracking components onto shadcn primitives Rebuilds the provider discount and margin tables, the pricing calculator and its multi-cost results on the in-repo shadcn layer, and swaps the imperative antd modal.confirm removals for AlertDialog. Row actions gained accessible names, which replace the Tremor stub mocks the tests used to drive. cost_tracking_settings keeps its two antd Modals and Forms, since they wrap the two add forms that stay on antd for now. * chore: retrigger e2e gate * feat(proxy): serve Anthropic-native /v1/models for Claude Code gateway discovery (#35455) * feat(proxy): serve Anthropic-native /v1/models for Claude Code gateway discovery * refactor(proxy): move Anthropic model-list formatter into llms/anthropic/common_utils * fix(proxy): make model_list request param optional for direct callers * style: apply ruff format to changed lines * style: satisfy ruff strict-rule budget (UP006, I001) * style: satisfy type-discipline budget (LIT002 mutable-ok, LIT009 pyright ignore) * style: satisfy LIT001/LIT010 and drop explanatory comment per contributor rules * fix(proxy): translate team model names in the Anthropic /v1/models response * ci: trigger buildkite status report * feat(proxy): carry token limits into the Anthropic-native /v1/models entries * fix(proxy): cast the injected request so the anthropic-version guard is a real comparison * fix(proxy): explain the model listing casts so the type-discipline gate passes --------- Co-authored-by: yuneng-jiang <[email protected]> Co-authored-by: Yassin Kortam <[email protected]> * fix(ui): keep the cost tracking removal confirmation open until it settles The discount and margin removal confirmation used AlertDialogAction, which renders AlertDialogPrimitive.Close and dismisses the dialog on click. The dialog therefore disappeared while the removal request was still in flight, leaving the admin with no sign that anything happened and free to fire a duplicate removal. Swap the confirm control for a plain destructive Button, track an isRemoving pending state that disables Cancel and relabels Remove to "Removing...", and clear the pending removal in a finally block once the request settles. * test(ui): build the deferred removal with Promise.withResolvers The pending-state test seeded its deferred promise by declaring the resolver with let and reassigning it inside the executor. Promise.withResolvers is the standard way to get the same handle without the reassignment, and the assertions are unchanged. * refactor(ui): declare DateRangePickerValue locally instead of importing it from tremor DateRangePickerValue is a plain object shape, not a component, so the twelve files that used it were each carrying a no-restricted-imports suppression for a type that tremor declares as { from?: Date; to?: Date; selectValue?: string }. Declare that shape in components/shared/date_picker_types.ts and point every consumer at it, which drops ten suppressions from the baseline. advanced_date_picker and usage_date_picker keep their tremor imports: they still render tremor Button, Text and DateRangePicker, and moving DateRangePicker itself needs react-day-picker. * refactor(ui): move MCP permission panels onto shadcn primitives Replaces antd Radio, Checkbox and Tooltip, plus Tremor Text and Badge, with the in-repo shadcn equivalents across the three MCP permission panels, and drops the no-restricted-imports suppressions they no longer need. Also removes the stale suppression on settings.test.tsx, which imports neither library. The tool rows keep their existing click-to-toggle behaviour: the row owns the toggle and the checkbox no longer carries its own change handler, since Base UI replays the click through a hidden input that reaches the row on its own. Adds payload-level tests for the risk-group view covering group clear, mixed-state re-arm, single-tool toggles from both the box and the row, and a controlled round trip proving each control re-renders from the permissions it emitted. * feat(proxy): add per-component response cost headers - Extract input_cost, output_cost, cache_read_cost, cache_creation_cost, reasoning_cost, and tool_usage_cost from logging object cost breakdown - Populate x-litellm-response-cost-* component headers in ProxyBaseLLMRequestProcessing.get_custom_headers - Ensure headers are omitted when cost breakdown is absent or values are None - Add comprehensive test suite covering component headers, math invariants, caching, reasoning, and discounts/margins * refactor(ui): migrate ten small dashboard files off antd and tremor Moves the onboarding views, router settings inputs, tag rate limit editor, fallback buttons, created-key display and the shared numerical input onto the in-repo shadcn layer. Each control has a direct equivalent, so this is a like-for-like swap with no layout changes and no new styling. Router settings saves by reading input values straight off the DOM with document.querySelector('input[name="..."]'), a path no test covered. Adds a regression test that types into a field and asserts the typed value reaches the payload, so the name attribute contract stays enforced. Also adds tests for TagRateLimitEditor, which had none and whose RPM cell switched from antd InputNumber to a native number input. * refactor(ui): drop explanatory comments from the migration tests * refactor(ui): drop narration comments from the MCP permission panels * fix(main): an explicit provider outranks a known OpenAI model name (#36800) * fix(main): an explicit provider outranks a known OpenAI model name completion() picks the OpenAI handler whenever `model in litellm.open_ai_chat_completion_models`, and that clause is evaluated before the gemini and vertex_ai branches. get_llm_provider() already resolves those names to "openai", so the clause only adds anything when the provider is something else, and then it silently overrides it: the config built for the requested provider is handed to the OpenAI handler. For gemini that is fatal. VertexGeminiConfig.transform_request raises NotImplementedError by design, since Vertex builds its request in its own handler, so `gemini/gpt-4o` dies in async_transform_request before anything is sent. register_model() reaches the same state without an odd model id: an entry claiming litellm_provider "openai" adds its name to open_ai_chat_completion_models, so one mislabelled pricing entry reroutes every later call to that model in the process. The name clause now applies only when no other provider was resolved. * test(main): move the routing regression into the mapped test file CLAUDE.md asks bug fixes to extend the mapped test file, so these belong in tests/test_litellm/test_main.py rather than a module of their own. They also no longer swap out the provider handler objects. Both Gemini cases inject an HTTPHandler whose post() answers like generativelanguage does, then assert the URL the request went to and read the reply back; the OpenAI case injects an OpenAI client and patches its own raw-response create. That asserts the endpoint the call reaches instead of which attribute the test replaced, and matches the neighbouring tests in the file. * fix(exception_mapping): bare 429 in an error body no longer outranks the status code (#36705) is_error_str_rate_limit treats any standalone 429 in the stringified exception as a rate limit, and for openai-compatible providers that check runs before the status-code branch. Providers echo the request back in validation errors, so a 400 whose body happens to contain a 429 comes out as RateLimitError. Tokenised prompts hit this routinely, since 429 is an ordinary token id (" that" in several tokenisers) and an echoed prompt_token_ids array is enough: {"error":{"message":"`tools` must not be an empty array", "type":"invalid_request_error","code":400}, "prompt_token_ids":[9906,429,1234]} The mislabel is not cosmetic. RateLimitError tells callers and routers to retry, so a request that cannot succeed gets replayed, and the failure is booked against provider throttling rather than the caller. Against DeepInfra, one recurring 400 ("`tools` must not be an empty array") came back as a rate limit in 77 of 198 occurrences, the split depending only on whether the echoed prompt contained 429. 16482 narrowed '"429" in error_str' to \b429\b after a false positive on 'asbjdad429addad'. Word boundaries cannot separate a real 429 from a token id, so the same class of false positive survives. is_error_str_rate_limit now takes an optional status_code, and the bare-number branch fires only when no explicit status contradicts it. The status is read off an arbitrary exception, so a non-integer is treated as unknown and left to the existing behaviour. The repo has a single call site. The phrase branches are untouched, so a provider reporting a real rate limit in the message text under a non-429 status still maps to RateLimitError (11455). This is not "status code wins". Tests cover the matcher (suppressed under a 400; still detected with no status, None, 429, or a non-integer status; phrase honoured under a 400) and exception_type end to end (400 with 429 in the echoed body -> BadRequestError, real 429 -> RateLimitError). Reverting the source change fails the latter. * fix(ui): distinguish hosted and local vLLM in the provider dropdown * test(vector_stores): drop redundant route-map comment * fix(proxy): force prisma recreate on postgres cached-plan error (#36428) `_query_first_with_cached_plan_fallback` recovers from Postgres's "cached plan must not change result type" by recreating the Prisma client, which drops both the server-side plans and the engine's client-side statement-name cache. Since #30183 the shared reconnect path probes the writer with `SELECT 1` first and skips the recreate when it answers, which is right for the IAM token refresh it was added for and wrong here: the connection is healthy, it is the session's prepared statements that are stale, so the probe always passes and always vetoes the recreate. Callers now pass `force_recreate` to skip that probe, and only the cached-plan fallback does. Getting past the probe is not enough on its own. Both cooldown checks would still skip the recreate for 15 seconds after any earlier reconnect, which outlives the 10 second auth retry window, so a migration landing in that window kept 503ing. `force=True` would fix that but would also let every concurrent caller of the same burst kill the engine the first one just built. The caller instead names the engine it observed before the query, and the cooldown is waived only while that engine is still the live one, so the first caller repairs the pool and the rest fall back to the normal cooldown. That engine has to be the one the query actually ran on. `query_first` is a top-level read, so with a read replica configured it is dispatched to the reader and it is the reader's prepared statements that go stale, while `writer_db` names a different engine with its own counter. The observation and the cooldown comparison both go through `read_db`, added alongside `writer_db` and backed by a `read_target` property on the routing wrapper that `__getattr__` now dispatches through so the two cannot drift. The observation carries the wrapper, not just its generation. `read_db` resolves to the reader while it is available and to the writer once it is not, and those counters are independent and both start at zero, so comparing a bare number across that switch pits one engine's counter against another's. Equal by coincidence waives the cooldown for an engine already replaced; unequal gates a caller that needs the recreate. Identity settles it, and is sound because the engine object is never re-pointed without the generation also moving. Three smaller holes on the way out. The waiver is withdrawn once a repair of that same engine has been tried and failed, so a burst collapses onto one attempt instead of each caller running its own recreate serially; the record is keyed per engine rather than counted globally, so an unrelated reconnect failure cannot suppress a stale reader's recovery and a writer failure cannot evict the reader's record. And a forced recreate that the optimistic-lock guard declines is no longer reported as a success on either the direct or the heavy path, since the routing wrapper leaves the reader untouched in that case; a decline is deliberately not counted as a failure, so the caller's own backoff still gets its waiver on the next attempt. A decline on the heavy path clears the dead-engine flag before raising. The clear after the cycle is skipped by any raise, which is right for a failure and wrong here, and the non-forced path already clears it on a decline, so this restores that policy rather than inventing one. Stranding the flag would route the next cycle back down the probe-free heavy branch, where the refreshed generation matches and the recreate kills the healthy engine a refresh just spawned, which is #29176. Clearing that flag is necessary and not sufficient. The escalation check re-arms it whenever the consecutive-failure count sits at the threshold, so a decline that left the count alone sent the very next attempt back down the same path. A decline is raised only at the generation guard, and the generation moves only after a replacement connects, so a decline is proof that a replacement succeeded and the count is reset on it. Fixes #36418 Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(transcription): stop a zero output rate from zeroing transcription cost (#36914) cost_per_second treated a declared-but-zero output_cost_per_second as a real rate, so the output branch claimed the call and the elif locked out input_cost_per_second. Every transcription model shipping output_cost_per_second 0.0 next to a real input rate billed $0, which covers 43 of the 55 per-second entries in the cost map: all 36 deepgram models, both assemblyai, both elevenlabs scribe, both groq whisper and azure-stt. Custom deployments pairing the two fields the same way billed $0 as well Take the output branch only when that rate is actually billable, so a zero falls through to the input rate. Entries that duplicate one rate into both fields, whisper-1 among them, keep billing exactly what they bill today * refactor(caching): accept read-only sequences for redis rpush pipeline payloads Keeps the spend buffer restore path free of mutable-collection construction so the type discipline gate stays within its LIT002 ceiling. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * Revert "fix(auth): stop the team fallback from widening model access (#36837)" This reverts commit ab2333b6c4d0fed2d78c55352a7c9d5aa53aea19. Every Admin UI login mints its session key against the sentinel team_id `litellm-dashboard`, and no LiteLLM_TeamTable row is ever created for it. That lookup is therefore a provably-absent row on every UI request, which #36837 turned into a hard refusal with no override, so the whole dashboard 404s. Reverting restores the token-derived fallback. The model-access widening #36837 closed is reopened and needs a re-land that exempts the UI sentinel team. * fix(langfuse): gate update_trace_keys behind an operator setting (#36862) update_trace_keys lets a caller name which request metadata entries get copied onto an existing trace, and the name is unrestricted. Sending update_trace_keys: ["user_api_key_auth"] with existing_trace_id serializes the resolved auth object, including the team callback credentials it carries, onto the trace through Langfuse.trace(**trace_params). TraceBody is Extra.allow, so an unexpected key ships rather than being dropped. Any holder of a team key can do this and read the result in the destination the team already logs to, so the feature is now inert unless an operator turns it on with langfuse_enable_update_trace_keys. * fix(fireworks_ai): let extra_body thinking/reasoning_effort take precedence over chat_template_kwargs * fix(ui): show zeroed auto-router usage stats when a window has no sessions (#36868) * fix(ui): show zeroed auto-router usage stats when a window has no sessions * test(ui): assert the muted track on the empty share-of-turns bar * fix(mcp): keep admin-entered oauth endpoints in management reads (#36888) * fix(mcp): keep admin-entered oauth endpoints in management reads Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(mcp): cover configured oauth endpoints on the config load path Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(ui): match the MCP servers count badge to its sibling permission badges The Object Permissions section rendered the MCP Servers badge with shadcn's default variant (solid bg-primary), so a plain count showed up as a black pill next to the light Vector Stores and Agents counts. Counts now use secondary everywhere, and destructive stays reserved for the blocked state. * refactor(ui): drop the explanatory comment from the badge variant test * fix(anthropic): bill undetailed iteration cache writes at the 5m rate * fix(cost_calculator): mirror the anthropic geo uplift in the token-type cost breakdown * fix(openai,azure): return a length-truncated 200 when the output budget fits no token (#36859) OpenAI and Azure GPT-5.x answer a chat request whose output budget cannot fit a single visible token with a 400, while the same models return a length-truncated 200 one or two tokens higher. Agents that probe a model with a hardcoded max_tokens of 1 read that 400 as "model unavailable". The four chat request helpers now recognise the provider's own sentence and hand back the length-truncated response the provider gives at a slightly larger budget: finish_reason "length", empty content, zero completion tokens. Any other 400 still raises. Streaming is covered by the same seam, and the caller's budget is never raised on their behalf. The provider bills the prompt it processed but sends no usage object with the 400, so the prompt tokens are estimated with the same token_counter every other usage-less path uses. Reporting zero would let a caller send an arbitrarily large prompt with max_tokens 1 and be charged nothing. * fix(proxy): always emit the Anthropic /v1/models token limits, null when unknown (#36961) Anthropic's Models API declares max_input_tokens and max_tokens as nullable, not optional, and the live vendor endpoint returns both keys on every entry. The merged Anthropic-native listing dropped either key whenever LiteLLM could not resolve a limit, so a client validating against a nullable-but-required schema saw a malformed entry for any model the cost map does not know. * feat(helm): add startupProbe and hpa.behavior to the componentized chart (#36382) Two small pod-spec passthroughs the componentized chart was missing, both additive and empty by default so existing renders are unchanged: - gateway/backend/ui deployments gain a `startupProbe` knob (same `{{- with }}` toYaml pattern as liveness/readiness), to gate liveness during a slow cold start without a kill loop. - gateway/backend/ui HPAs gain an `hpa.behavior` passthrough rendered verbatim under spec.behavior (scaleUp/scaleDown policies + stabilization windows). Tests: extend probe_tests.yaml (startupProbe absent by default / renders verbatim) and add hpa_behavior_tests.yaml. Full chart suite: 76 tests pass. Signed-off-by: Louis Vauterin <[email protected]> Co-authored-by: Claude Opus 4.8 <[email protected]> * fix(vector_stores): classify write endpoints before reads on substring collisions * fix(router): stop get_router_model_info from wiping cached pricing Merge deployment model_info into a copy of the lru_cache'd get_model_info() dict and drop unset Nones, so Deployment's mirrored pricing defaults no longer overwrite built-in prices process-wide. Fixes #36980 Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(bedrock): resolve aliases in batch file records * fix(caching): tolerate SSE chunk splits in anthropic stream cache writer * fix(proxy): return cost breakdown header values as a named tuple * fix(responses_api): map bridged chat usage on guardrail-blocked replies Move the blocked-usage mapping for /v1/responses next to blocked_response_usage in guardrail_translation utils, map bridged chat prompt/completion tokens to Responses API input/output tokens, and let raise_passthrough_exception attach the blocked response so post-call guardrail blocks report real usage * fix(proxy): serve aggregate MCP endpoint on bare /mcp instead of 307-redirecting (#34845) The MCP sub-app is attached with app.mount("/mcp", ...) and a Starlette mount never matches its bare prefix, so POST /mcp fell through to the router's redirect_slashes 307. Behind a TLS-terminating ingress whose peer address is not in uvicorn's forwarded-allow-ips (default: loopback only) the redirect Location is built from the socket scheme as http://, and MCP clients strip the Authorization header on the cross-origin follow, so reconnects fail with ECONNRESET right after a successful OAuth flow. The redirect also fires before auth, so the bare spelling never returns the RFC 9728 WWW-Authenticate challenge that OAuth clients need to start the flow. Add an explicit /mcp route beside the existing /toolset/{name}/mcp and /{name}/mcp spellings, forwarding to handle_streamable_http_mcp with the same scope rewrite those routes already use (path=/mcp, _original_path preserved for OAuth challenge URL selection). When the mcp package is unavailable the route 404s, matching what the bare sub-app serves on /mcp/ in that state. /mcp/, /mcp/{server}, /{server}/mcp and /toolset/{name}/mcp spellings are unchanged; the exact-match route and the mount have disjoint match sets so registration order cannot matter. * fix(cost): reach tiered pricing for models without top-level per-token rates * feat(shadow_eval): add reverse-direction shadow eval jobs (#36865) Shadow eval only answered "should this key adopt this auto-router". Once a key is on the router it is invisible to the feature, because the sampling gate skips any request the shadowed router already served, so post-adoption quality regressions go unmeasured. Reverse mode inverts the arms: sample the traffic the router did serve and duplicate it against a fixed baseline_model, judged by the same blind pairwise judge. Same job table, same attempt rows, same aggregates. real_* stays the arm the caller was served and shadow_* the duplicated one, so in reverse real_model is the router's pick and shadow_model is the baseline. The active-job slot becomes one per (key, direction) so both directions can run at once, and tier attribution in reverse reads the control request's routing decision rather than the shadow call's write-back. * fix(fireworks_ai): move top-level thinking into extra_body on the text completion path * feat(search): add Nimble as a search provider (#36347) * feat(search): add Nimble as a search provider Adds `NimbleSearchConfig` so `search_provider: nimble` works across the SDK, the proxy /v1/search endpoint, the Search Tools dashboard, and spend tracking. Nimble's /v2/search already uses the Perplexity unified spec's parameter names, so the request transform is close to a pass-through. `search_domain_filter` splits into include_domains/exclude_domains on the spec's `-` prefix, `country` is upper-cased to the ISO form Nimble documents, and everything else is forwarded so focus, search_depth, time_range and the rest stay reachable. On the response side, snippet prefers `content` and falls back to `description`, and a malformed body raises an attributed error rather than reporting an empty search. Also tightens `BaseSearchConfig.get_supported_perplexity_optional_params` to return `frozenset[str]` instead of a bare mutable `set`, which every caller already treats as read-only. * fix(search): surface Nimble error bodies instead of empty results Greptile flagged that a null or absent `results` degraded to a successful empty search. A search with no hits comes back as `"results": []`, verified against the live API, so the field is now required and anything else raises the attributed schema error the other malformed bodies already take. Also unwraps Nimble's second error envelope. Collection failures return `{"success", "task_id", "message"}` rather than the `{"detail"}` shape validation errors use, and only the latter was being read. Drops comments that restated the adjacent code. * docs(search): drop the Nimble param list from the transform docstring It restated the vendor's API reference, which the module docstring already links, and would go stale the moment Nimble adds a focus mode. * fix(bedrock): fall back to the batch deployment model for unmapped record models * fix(cost-tracking): count dict-shaped web_search_call output items * fix(mcp): drop caller host and configured upstream headers from logged metadata (#36901) * fix(mcp): drop caller host and configured upstream headers from logged metadata The synthetic request that carries MCP client headers into add_litellm_data_to_request forwarded the caller's Host header, and Request.url is built from it, so a caller chose the proxy_server_request url and the metadata endpoint that every logging callback records. _upstream_credential_headers also only knew the configured client side auth header and the x-mcp- prefix family, so a header name declared in mcp_servers.<name>.extra_headers reached logging metadata in cleartext. Those names are admin chosen, so no prefix rule can recognize them; read them off the server registry instead. The header is still forwarded upstream, which is what extra_headers is for. authorization is left out because clean_headers already strips it and claiming it here would move authenticated_with_header on the oauth passthrough config. The Responses bridge tests stub the server manager, so their fakes gain the registry accessor the sanitizer now reads. * fix(mcp): drop caller host from the sanitized header mapping too The synthetic request stopped forwarding host, but the parallel sanitizer did not, so a forged hostname still reached the guardrail payload and the list_tools spend row. Drop it there as well. Exempt the configured identity headers from the upstream credential set. get_user_from_headers resolves end user attribution off the same request this module reconstructs, and it only fills end_user_id when auth left it unset, so claiming user_header_name or a user_header_mappings name would lose attribution on the MCP paths that authenticate upstream. Drop the isinstance guard on extra_headers entries: the field is typed list[str], so the check is dead and basedpyright scores it. * fix(mcp): accept a bare user_header_mappings entry when exempting identity headers get_internal_user_header_from_mapping and get_customer_user_header_from_mapping both normalize a single mapping to a one element list, and config_settings.md documents the key as a dict. Iterating the bare form yields its keys instead, so the exemption silently matched nothing and an identity header also named in an MCP server's extra_headers was dropped after all. * fix(router): honor tiered_pricing set in a deployment's litellm_params * feat(scripts): queue heavy gates behind a machine-wide slot lock * test(proxy): assert production nesting semantics for component cost headers * fix(cost): bill reasoning tokens at the selected tier's reasoning rate * fix(vertex_ai): fail an embeddings batch entry whose fan-out came back incomplete Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(router): merge model_info without new mutable constructions Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(cost-tracking): price web search on dated search-preview map entries * fix(proxy): emit uncached input cost so component headers sum to the total * refactor(ui): re-sync badge and skeleton onto the base-vega shadcn style components.json has declared "style": "base-vega" since cfe9e39e55, but badge and skeleton were added a few days earlier under new-york and never re-synced, so both still carried the previous style's classes. Badge's destructive variant rendered as solid red with white text instead of the tinted wash the rest of the dashboard uses, which is already the convention for Button Re-runs npx shadcn add for both and keeps the two local deltas the registry cannot supply: cva comes from @/lib/cva.config, since class-variance-authority is not a dependency here, and both stay wrapped in React.forwardRef, which the tripwire in tests/setupTests.ts requires until the React 19 upgrade Adds Badge to ref-forwarding.test.tsx. Nothing covered it before, even though two TooltipTrigger sites compose over it, so the wrapper could have been dropped by the next re-sync without a single test going red Retargets one assertion in LogDetailContent.test.tsx. It regex-matched the whole class string for "destructive" to prove a tag was not alarming red, which the restored aria-invalid classes now satisfy for every variant; it checks the variant attribute and red utility classes instead * test(vertex_ai): cover duplicated fan-out rows in embeddings batch reassembly Also ruff-formats the batch transformation test file, which the formatter gate flags once the file is touched. * test(ui): assert cache and retry tags by text instead of class name Three assertions in LogDetailContent.test.tsx matched a regex against the rendered class string to prove a tag was green or was not red. That pins styling rather than behavior, and jsdom does not resolve the utilities anyway, so the checks only ever proved that a substring survived into the class attribute The badge re-sync exposed it: base-vega's base string carries aria-invalid variants of the destructive token, so a "not destructive" regex started matching every badge regardless of variant Each one now asserts the tag's text is present, which is what the surrounding cases already do and what the user actually observes * fix(ui): stop the models tab strip from scrolling vertically The tab strip carried overflow-x-auto directly on the TabsList. CSS forces overflow-y from visible to auto once overflow-x is not visible, and the line variant's active-tab underline is an absolutely positioned ::after that hangs 5px below its trigger, so the strip picked up a pixel of vertical scroll on top of the horizontal scroll it actually wants. The scroll container now lives on a wrapper whose bottom padding leaves room for the underline, offset by a matching negative margin so the row keeps its exact geometry. * fix(anthropic_messages): make tool_result images visible to OpenAI-compatible providers (#34462) Images nested inside an Anthropic `tool_result` block were dropped when the request was adapted for an OpenAI-compatible provider, because the OpenAI tool message shape only carried text. Hoist those images out of the tool result and into a following user message so the model can still see them, and widen the tool message content type to accept image parts. * fix(ui): anchor chips-combobox popups to the field instead of the inner input Base UI positions a combobox popup against the Combobox.Input by default. In chips mode the visible field is the ComboboxChips wrapper and the input is a smaller box nested inside it, so every chips-combobox in the dashboard opened its popup 11px right of the field and 17px past its right edge. shadcn ships the wiring for this and their combobox-multiple example uses it: useComboboxAnchor on the chips container, passed to ComboboxContent as anchor. The anchor prop also drives data-chips, which cancels the extra min-width an ordinary combobox wants. Every chips site in the dashboard omitted it. The anchor is attached through Base UI's render prop rather than a plain ref, because React 18 drops refs on function components and ComboboxChips is one. Adds MultiSelect's first test, covering the anchor wiring plus selection, chip rendering and custom values. * test(ui): assert which element the chips-combobox popup anchors to The previous assertion read data-chips, which is derived from the anchor prop being truthy, so it stayed true even when the ref never reached the DOM and the popup was still anchored to the inner input. Stub distinct widths on the chips container and the input, then read the width the positioner resolved. Reverting the anchor wiring now reports the input's width instead of the field's, which is the actual bug. * refactor(cost): make the shared token-details parsers public parse_prompt_tokens_details and parse_completion_tokens_details are imported by four modules, so the leading underscore made every import a reportPrivateUsage violation * fix(ptu): clear a PTU deployment's tiered_pricing instead of zeroing it tiered_pricing is a list, so the 0.0 the flat-rate zeroing stores does not even validate. Supplying tiers alongside PTU config gets the same 400 as a flat rate; tiers already stored are dropped from both blobs * fix(ptu): empty a PTU deployment's tiered_pricing instead of dropping it Dropping it falls back to the public cost map's tier table, whose rates outrank the zeros written beside them, so a PTU deployment on a tiered model keeps billing its traffic per token. Stored empty, the tiers no longer apply and the zeros win * fix(cost): fall back to the model output rate when a tier omits one Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(cost): inherit the backend output rate when a deployment's tiers omit one * fix(dashscope): honor the model reasoning rate when a tier omits output rates Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): requeue spend logs when the DB write fails with a transport error (#36716) * fix(proxy): requeue spend logs when the DB write fails with a transport error Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): hardcode the spend log queue cap and drop the stale re-export Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(proxy): keep the spend log requeue within the type discipline budget Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): apply the spend log queue cap to producer appends too Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): lower the spend log queue cap to 1k and make it env configurable Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): bound the spend log queue by bytes instead of row count A row cap cannot bound memory: a row carries the whole prompt under store_prompts_in_spend_logs, so a cap that rides out an outage of counter-only rows is an OOM once prompts are stored. Every enqueue and dequeue now goes through one pair that tracks what the queue costs and drops the oldest rows past a 64 MB budget. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): make the spend log queue byte budget env configurable Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): use a string default …
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TLDR
Problem this solves:
How it solves it:
User Flow
Before: a developer transcribing audio through the proxy is billed nothing, so spend tracking and budgets never see the usage
"model": "deepgram/nova-3"and a 60 second audio file"spend": 0.0After: the same request records real spend, so logs and budgets reflect the audio actually transcribed
"model": "deepgram/nova-3"and the same 60 second file"spend": 0.0043for that entryRelevant issues
Fixes #19154
Linear ticket
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Live proxy, real provider APIs, no mocks. Same config, same audio file (
tests/gettysburg.wav, 17.58s), same two curls on both commits.deepgram/nova-3is one of the 43 affected models;whisper-1is the unaffected control that must not changeBefore, at f03df1b (latest litellm_internal_staging): the deepgram transcription succeeds but bills nothing, no cost header is returned at all
After, at 24406d6 (this PR): the same deepgram request now returns the real cost, and the control is byte-for-byte unchanged
Sanity check: 17.577s of audio at deepgram/nova-3's
input_cost_per_secondof 7.167e-05 is 0.00125975, matching the returned header exactly. whisper-1 bills 17.57s at 0.0001/s on both commits, confirming the duplicate-rate entries keep billing exactly what they bill todayType
🐛 Bug Fix
✅ Test
Caveats (if any)