feat(proxy): add paginated GET /public/v1/model_hub - #38636
Conversation
…package The list framework and its RFC 9457 problem machinery sat under management_endpoints/management_v1/, which was the right home while /management/v1 was its only consumer. The public surface is about to build on the same framework, and a control-plane package is the wrong thing for a public route to import. Moves list_framework.py in full, plus everything in common.py except MANAGEMENT_V1_PREFIX, to litellm/proxy/list_api/. Every importer is updated directly instead of leaving re-export shims, so each symbol keeps exactly one import path. ManagementProblem keeps its name: renaming it would touch the app-wide exception handler and every call site for no behavioural gain. The framework's own tests move alongside the code they cover. The fastapi removed-name guard in test_common.py now globs both packages, so budgets.py and spend_logs.py stay covered after leaving the framework's directory. Pure move, no behaviour change: the 179 tests across both packages pass unchanged.
The public Model Hub page loads every public model group in one call.
Measured on a live proxy with 300 published groups, /public/model_hub
answers with 328 KB in a single response and the page renders all 300 rows
into the DOM. At a few thousand models that is multiple megabytes and a
page that stops responding, which is what a customer reported.
Adds GET /public/v1/model_hub, the first resource on the unauthenticated
/public/v1 surface. It is built on the shared list framework, so it gets
the {data, meta, links} envelope, RFC 9457 problems, strict unknown and
duplicate query parameter rejection, and sort validation without
reimplementing any of it. Sorting covers model_group, mode, the token
limits and the per-token costs, `q` searches model_group, and the filters
are the ones the page actually offers: mode and providers. Default sort is
alphabetical, which is what a browse list wants and what these rows can
support: they carry no creation timestamp.
/public/model_hub is untouched. The shipped UI still calls it and its
migration is a separate change, so this is purely additive alongside it.
Model hub rows are computed off the running router rather than read from a
table, so this adds InMemoryListExecutor: the same QueryPlan applied in
Python instead of rendered to SQL. It matches the SQL executors where it
counts, NULLS LAST in both sort directions and NULL satisfying no
comparison, so a filter means the same thing on either. The other three
public hubs have the same shape and can reuse it as is.
The fix itself is ordering. The endpoint being superseded reads every
latest health check and joins it against the whole model list, so paging
the response alone would have changed nothing. Here the health lookup is
an injected dependency the executor calls on the page slice, after the
filter and the sort, so it resolves health for the rows being served and
no others. PrismaClient gains a bounded read for that, next to the
unbounded one it mirrors. The regression test pins the ordering by
asserting which model groups the lookup is asked about, and fails against
an enrich-then-slice implementation.
Five adversarial review passes over the branch. What they found: `is_null` was the one predicate in the in-memory executor that read a repeated field's container instead of its elements, so a field holding only nulls was indistinguishable from a populated one. It now lifts over elements like every other predicate does. Not reachable through this endpoint, whose only repeated field grants `contains` alone, but the executor is written to be reused by the other three hubs and the inconsistency was a trap for them. The fastapi removed-name guard globbed the framework packages but not `public_endpoints/public_v1`, which `proxy_server` also imports unguarded at module level, so the new package had none of the protection the test claims to give. It now covers all three. Regenerates the dashboard's API types, which the OpenAPI sync check requires whenever the proxy's route surface moves. The diff is the 65 generated lines for the new operation and nothing else; no dashboard code changes here. Also trims comments and docstrings that argued for a decision or restated a signature rather than explaining code, and wraps a docstring line that ran past 120 characters.
The framework's tests moved from tests/test_litellm/proxy/management_endpoints, which the proxy-endpoints shard claims, into a new tests/test_litellm/proxy/list_api that no shard named. Both coverage guards caught it: the semantic shards have no catch-all bucket, so the directory would have run nowhere. Claims it alongside management_endpoints, where the same tests ran before.
Greptile SummaryThe PR adds a paginated public Model Hub endpoint while retaining the existing unpaginated route
Confidence Score: 5/5The PR appears safe to merge No blocking failure remains
|
| Filename | Overview |
|---|---|
| litellm/proxy/public_endpoints/public_v1/model_hub.py | Defines the paginated public catalogue, page-bounded health enrichment, and public response contract; both prior review concerns are resolved or withdrawn |
| litellm/proxy/list_api/in_memory.py | Implements in-memory predicate evaluation, stable multi-key ordering, page slicing, and post-slice enrichment |
| litellm/proxy/list_api/list_framework.py | Moves the shared list framework to a surface-neutral package and rejects duplicate sort fields |
| litellm/proxy/list_api/common.py | Centralizes list-route problem responses, query validation, escaping, and pagination links |
| litellm/proxy/utils.py | Adds a bounded health-check lookup for only the model groups present on the requested page |
| litellm/proxy/proxy_server.py | Registers the new public endpoint router and shared problem-response handling |
| litellm/proxy/_types.py | Classifies the new Model Hub route as public so invalid or absent credentials do not block it |
| ui/litellm-dashboard/src/lib/http/schema.d.ts | Regenerates the API schema declarations for the new paginated endpoint |
Reviews (3): Last reviewed commit: "fix(proxy): clear the two basedpyright e..." | Re-trigger Greptile
PR overviewAll previously flagged issues have been addressed. No open security concerns remain on this pull request. Security reviewNo open security issues remain on this pull request. Fixed/addressed: 1 · PR risk: 0/10 |
…tring The docstring listed every sortable field, the page-size cap and the filter set, all of which already live in MODEL_HUB_LIST_SPEC and all of which the endpoint hands back in the allowed array of a rejected request. Two copies of one spec is a prose update owed on every change to the real one. Keeps what a caller cannot derive from the endpoint itself: what the resource is, that it needs no authentication, and a working example. Regenerates the dashboard types, which carry the docstring as the operation description.
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
sort took any number of comma-separated keys, and the in-memory executor runs one full sorted() pass per key before slicing. Naming one allowed field N times therefore bought N passes over every published model group, synchronously on the event loop, from a route that needs no credentials. Measured on 300 groups: 0.001s for one key, 0.034s for a thousand, 0.166s for five thousand, and it grows with the catalogue this endpoint exists to make large. A repeated field cannot change the ordering, so rejecting repeats costs a caller nothing and bounds the passes at len(sortable), a number the spec author picks rather than the caller. That beats an arbitrary cap: no magic number, and the bound holds for every resource built on the framework. The tiebreaker is appended after parsing, so sorting explicitly by it stays legal. Budgets renders one ORDER BY in SQL and never had the amplification, but the check belongs with the rest of the sort validation rather than in one executor.
|
@greptileai review Generated by Claude Code |
Two CI gates, one cause. AnyOf declared its clauses as Predicate, so both consumers had to recurse to evaluate one: the SQL renderer through _render/_render_all, and the in-memory executor through _holds. The recursion detector flags the latter, and its reason is the same one this PR already ran into once, a caller-controlled cost that shows up as CPU. Nothing actually builds a nested AnyOf. _search_predicate is its only producer anywhere in the repo and it emits Compare leaves, in every call site and every test. Declaring clauses as tuple[Compare, ...] makes that a fact the type checker keeps rather than a comment, and _holds then evaluates a disjunction of leaves with no recursion at all. Also marks the new health read's broad except, which the strict gate counts, and covers the ordering comparison operators. The endpoint exposes only eq/in/contains, so gt/gte/lt/lte were live code no test evaluated.
The bounded health query added ten LIT002 violations, which pushed the codebase total past its budget. The gate counts across the tree and compares to the merge base, so a file already carrying debt does not absorb new violations. Returns an empty tuple rather than an empty list on the two no-result paths: the signature already promises a Sequence, so that is a free two-violation reduction and a better type. Builds prisma's order argument from a tuple of pairs, which turns four literals into one. The three that remain are prisma's own API shape and each carries its reason. Both budget gates now pass against the merge base.
The type-check budget is over its ceiling on the base already, so the gate blames any increase: reportArgumentType 2574/2564 and reportPrivateUsage 1815/1808, one each, both from this file. fastapi types a route's tags as list[str | Enum], so the tuple was an argument error; budgets.py has the same one and it is part of what put the rule over. Passing a list is what the signature asks for, marked because an inline list is a construction the discipline gate counts. _get_model_group_info is private by name but is the shared reader the endpoint this supersedes imports the same way, so the import carries a rule-scoped ignore with that reason rather than a copy of the function. basedpyright now reports zero errors across both new modules, and all three budget gates pass against the merge base.
|
@greptileai review Generated by Claude Code |
|
QA passed at |
TLDR
Problem this solves:
How it solves it:
GET /public/v1/model_hub, first route on/public/v1/public/model_hubis untouched, so the shipped UI keeps workingUser Flow
Before: someone browsing a proxy's public model catalogue waits on a single huge response, and on a large proxy the page stops responding
After: the same person asks for one page at a time, and gets an answer sized to what they can actually see
{data, meta, links};metasaystotal_count: 300andtotal_pages: 12links.nextis a ready-made URL for page 2, andlinks.lastfor page 12, so paging forward is one request?q=claude, sort with?sort=-input_cost_per_token, or filter with?filter[providers][contains]=anthropic, and each answer is one page?limit=10, comes back as a 400 that names every parameter it does accept, instead of being silently ignoredRelevant issues
Linear ticket
Resolves LIT-5410
Pre-Submission checklist
uv run pytest tests/test_litellm/proxy/public_endpoints/public_v1/test_model_hub.py tests/test_litellm/proxy/list_api/ -vScreenshots / Proof of Fix
Shared setup, used identically for Before and After: a config publishing 300 public model groups across 5 providers and 3 modes, a real Postgres with one health-check row per group, and the proxy booted with
Postgres statement logging is on, so the health-check read each endpoint issues is quoted from the database's own log.
Before (39e5b0c)
Case 1: there is no paginated public model hub endpoint
curl -s -o /dev/null -w "%{http_code}\n" "http://localhost:4000/public/v1/model_hub?page=1&page_size=25"curl -s "http://localhost:4000/public/v1/model_hub?page=1&page_size=25"Case 2: the only public model hub endpoint returns every group in one response
curl -s http://localhost:4000/public/model_hub | wc -ccurl -s http://localhost:4000/public/model_hub | jq lengthCase 3: it reads every health check row, not just the ones it shows
Case 4: sort, search, filter and the error contract
Case 5: a repeated sort key
After (04cbe79)
Case 1: the paginated endpoint serves one page
curl -s "http://localhost:4000/public/v1/model_hub?page=1&page_size=25" | jq -c '{meta, links, rows: (.data|length)}'{"meta":{"total_count":300,"page":1,"page_size":25,"total_pages":12},"links":{"self":"/public/v1/model_hub?page_size=25&page=1","first":"/public/v1/model_hub?page_size=25&page=1","prev":null,"next":"/public/v1/model_hub?page_size=25&page=2","last":"/public/v1/model_hub?page_size=25&page=12"},"rows":25}Case 2: the same 300 groups, one page instead of the whole list
curl -s http://localhost:4000/public/model_hub | wc -ccurl -s "http://localhost:4000/public/v1/model_hub?page=1&page_size=25" | wc -ccurl -s http://localhost:4000/public/model_hub | jq -c '{type: type, length: length}'(the old endpoint still answers with its bare array){"type":"array","length":300}Case 3: it reads health only for the rows it shows
"http://localhost:4000/public/v1/model_hub?page=1&page_size=10", then read the health-check query out of the Postgres log:http://localhost:4000/public/model_hubon the same build and read its query, to confirm the old endpoint is unchanged:curl -s "http://localhost:4000/public/v1/model_hub?page_size=2" | jq -c '[.data[]|{model_group,health_status,health_response_time}]'[{"model_group":"anthropic-model-001","health_status":"healthy","health_response_time":101.0},{"model_group":"anthropic-model-006","health_status":"unhealthy","health_response_time":106.0}]Case 4: sort, search, filter and the error contract
curl -s "http://localhost:4000/public/v1/model_hub?sort=-model_group&page_size=5" | jq -c '[.data[].model_group]'curl -s --globoff "http://localhost:4000/public/v1/model_hub?filter[providers][contains]=mistral&page_size=3" | jq -c '{total_count:.meta.total_count, rows:[.data[]|{model_group,providers}]}'{"total_count":60,"rows":[{"model_group":"mistral-model-002","providers":["mistral"]},{"model_group":"mistral-model-007","providers":["mistral"]},{"model_group":"mistral-model-012","providers":["mistral"]}]}curl -s --globoff "http://localhost:4000/public/v1/model_hub?filter[mode][in]=embedding,image_generation&page_size=3" | jq -c '{total_count:.meta.total_count, rows:[.data[]|{model_group,mode}]}'{"total_count":60,"rows":[{"model_group":"deepseek-model-009","mode":"image_generation"},{"model_group":"deepseek-model-019","mode":"image_generation"},{"model_group":"deepseek-model-029","mode":"image_generation"}]}curl -s "http://localhost:4000/public/v1/model_hub?q=model-04&page_size=5" | jq -c '{total_count:.meta.total_count, rows:[.data[].model_group]}'{"total_count":10,"rows":["anthropic-model-041","anthropic-model-046","deepseek-model-044","deepseek-model-049","groq-model-043"]}curl -s "http://localhost:4000/public/v1/model_hub?limit=10"(unknown parameter, HTTP 400,application/problem+json){"type":"urn:litellm:error:unknown-query-parameter","title":"Unknown query parameter","status":400,"detail":"Unrecognized query parameter(s): limit.","allowed":["filter[mode]","filter[mode][in]","filter[providers][contains]","page","page_size","q","sort"]}curl -s "http://localhost:4000/public/v1/model_hub?sort=providers"(undeclared sort field, HTTP 400){"type":"urn:litellm:error:invalid-sort-field","title":"Invalid sort field","status":400,"detail":"Cannot sort model groups by: 'providers'.","allowed":["input_cost_per_token","max_input_tokens","max_output_tokens","mode","model_group","output_cost_per_token"]}curl -s -o /dev/null -w '%{http_code}\n' -H 'Authorization: Bearer sk-not-a-real-key' "http://localhost:4000/public/v1/model_hub?page_size=2"(a wrong key must not make a public route 401)Case 5: a repeated sort key is refused rather than sorted twice
curl -s "http://localhost:4000/public/v1/model_hub?sort=model_group,model_group"(HTTP 400,application/problem+json){"type":"urn:litellm:error:duplicate-sort-field","title":"Duplicate sort field","status":400,"detail":"Sort field(s) named more than once: model_group. Each may appear once.","allowed":["input_cost_per_token","max_input_tokens","max_output_tokens","mode","model_group","output_cost_per_token"]}What the adversarial review sweep covered beyond the committed tests
Five review passes ran over the branch before it was opened. Beyond what is pinned in the test files, they exercised and found clean:
ORDER BY ... NULLS LASTcomparator: zero mismatches across mixed ascending and descending keys, nulls, string, int and float fields, and tiebreaker tieslinks.nextalone at page sizes 7, 13, 17 and 50: 300 unique rows, no gaps and no duplicates on any walk, with the correct remainder on every last page/public/model_hubagainst a full walk of the new endpoint: byte-identicalq=model_001matching nothing whileq=model-001matches one, proving_and%are escaped as literals rather than treated as wildcards; a quote-and-semicolon search string returns nothing and leaves the table intactpage_sizeof 101 and 100000 clamping to 100;pageandpage_sizegiven as negatives, floats, hex, text and a 24-digit integer all answered rather than crashingx-api-keyall returning 200, never 401/publicthat could shadow the new oneType
🆕 New Feature
Caveats (if any)
Medium
filter[providers][contains]matches a substring, not a whole provider name/management/v1/budgetstooLow
ui/litellm-dashboard/diff is the generatedschema.d.tslines only/public/model_hubsupports_*are booleans andFilterSpechas no boolean typeManagementProblemkeeps its name in the new neutral packagepyproject.toml's mutmut scope still lists onlymanagement_endpoints/pyproject.tomldiff emptyFinal Attestation