feat(a2a): semantic search over the agent registry via GET /v1/agents?query and an agent_search MCP tool - #38609
Conversation
…?query and an agent_search MCP tool
…/ tools/call validates
Greptile SummaryAdds semantic ranking to the accessible A2A agent registry and exposes it through REST and MCP while attributing embedding spend to the calling key.
Confidence Score: 5/5The PR appears safe to merge because no blocking failure remains. No blocking failure remains.
|
| Filename | Overview |
|---|---|
| litellm/proxy/agent_endpoints/agent_search.py | Implements embedding-backed ranking, spend metadata propagation, dimension recovery, and concurrency-safe cache merging. |
| litellm/proxy/agent_endpoints/endpoints.py | Adds validated search parameters and maps search configuration or embedding failures to explicit HTTP responses. |
| litellm/proxy/agent_endpoints/auth/agent_permission_handler.py | Extracts the existing accessible-agent filtering logic for reuse by listing and search. |
| litellm/proxy/_experimental/mcp_server/tool_search.py | Defines and dispatches the agent_search virtual tool through the shared agent-search implementation. |
| litellm/proxy/_experimental/mcp_server/rest_endpoints.py | Routes agent_search calls from the MCP REST endpoint through the virtual-tool path. |
| litellm/proxy/_experimental/mcp_server/server.py | Exposes and dispatches agent_search through the protocol MCP server with the existing virtual-tool permission gate. |
| tests/test_litellm/proxy/agent_endpoints/test_agent_search.py | Covers ranking, cache concurrency, mixed vector dimensions, permission filtering, spend metadata, and error outcomes. |
| tests/test_litellm/proxy/_experimental/mcp_server/test_mcp_tool_search.py | Covers virtual-tool schema serialization, discovery, authorization, and agent-search dispatch. |
Reviews (5): Last reviewed commit: "fix(a2a): merge fresh agent vectors into..." | Re-trigger Greptile
Codecov Report❌ Patch coverage is 📢 Thoughts on this report? Let us know! |
PR overviewAll previously flagged issues have been addressed. No open security concerns remain on this pull request. Security reviewNo open security issues remain on this pull request. Fixed/addressed: 1 · PR risk: 0/10 |
|
bugbot run |
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
Autofix Details
Bugbot Autofix prepared a fix for the issue found in the latest run.
- ✅ Fixed: Stale embeddings crash agent ranking
- After embedding, the index now drops cached agent vectors whose dimension no longer matches the query and re-embeds those texts, so a dimension change from a router fallback or config reload can no longer poison the cache with mismatched vectors that make cosine_similarity raise ValueError.
Or push these changes by commenting:
@cursor push 2ed2486bf4
Preview (2ed2486bf4)
diff --git a/litellm/proxy/agent_endpoints/agent_search.py b/litellm/proxy/agent_endpoints/agent_search.py
--- a/litellm/proxy/agent_endpoints/agent_search.py
+++ b/litellm/proxy/agent_endpoints/agent_search.py
@@ -164,10 +164,24 @@
return AgentSearchEmbeddingFailed(
reason=f"embedding model returned {len(vectors)} vectors for {len(missing) + 1} inputs"
)
- self._vectors = MappingProxyType(dict(chain(self._vectors.items(), zip(missing, vectors[1:], strict=True))))
+ query_vector: Final = vectors[0]
+ fresh: Final = dict(zip(missing, vectors[1:], strict=True))
+ kept: Final = {text: vec for text, vec in self._vectors.items() if len(vec) == len(query_vector)}
+ stale: Final = tuple(dict.fromkeys(text for text in texts if text not in kept and text not in fresh))
+ try:
+ refreshed: Final = await embed(stale) if stale else ()
+ except (OpenAIError, ValueError, BudgetExceededError) as exc:
+ return AgentSearchEmbeddingFailed(reason=f"re-embedding stale agent texts failed: {exc}")
+ if len(refreshed) != len(stale):
+ return AgentSearchEmbeddingFailed(
+ reason=f"embedding model returned {len(refreshed)} vectors for {len(stale)} inputs"
+ )
+ self._vectors = MappingProxyType(
+ dict(chain(kept.items(), fresh.items(), zip(stale, refreshed, strict=True)))
+ )
ranked: Final = sorted(
(
- AgentSearchHit(agent=agent, score=cosine_similarity(vectors[0], self._vectors[text]))
+ AgentSearchHit(agent=agent, score=cosine_similarity(query_vector, self._vectors[text]))
for agent, text in zip(agents, texts, strict=True)
),
key=lambda hit: hit.score,
diff --git a/tests/test_litellm/proxy/agent_endpoints/test_agent_search.py b/tests/test_litellm/proxy/agent_endpoints/test_agent_search.py
--- a/tests/test_litellm/proxy/agent_endpoints/test_agent_search.py
+++ b/tests/test_litellm/proxy/agent_endpoints/test_agent_search.py
@@ -147,7 +147,20 @@
outcome = await AgentSearchIndex().search("q", AGENTS, top_k=5, embed=short)
assert isinstance(outcome, AgentSearchEmbeddingFailed)
+ @pytest.mark.asyncio
+ async def test_dimension_change_invalidates_cached_agent_vectors(self) -> None:
+ index = AgentSearchIndex()
+ warm = FakeEmbedder()
+ await index.search("language translation", AGENTS, top_k=5, embed=warm)
+ async def wider(texts: Sequence[str]) -> Sequence[Vector]:
+ return tuple((1.0, 0.0, 0.0, 0.0) for _ in texts)
+
+ outcome = await index.search("language translation", AGENTS, top_k=5, embed=wider)
+ assert isinstance(outcome, AgentSearchHits)
+ assert len(outcome.hits) == len(AGENTS)
+
+
class TestSearchAgents:
@pytest.mark.asyncio
async def test_no_embedding_model_is_not_configured(self) -> None:You can send follow-ups to the cloud agent here.
…-embed on dimension changes
…vectors change dimension
|
bugbot run |
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
Autofix Details
Bugbot Autofix prepared a fix for the issue found in the latest run.
- ✅ Fixed: Cache keeps mixed-dimension agent vectors
- Rebuilt the cache write to re-read the current per-model map and filter both it and the embed result down to entries matching the new query dimension, so a subset re-embed evicts old-size vectors instead of merging them back from a stale pre-await snapshot.
Or push these changes by commenting:
@cursor push efbeae4393
Preview (efbeae4393)
diff --git a/litellm/proxy/agent_endpoints/agent_search.py b/litellm/proxy/agent_endpoints/agent_search.py
--- a/litellm/proxy/agent_endpoints/agent_search.py
+++ b/litellm/proxy/agent_endpoints/agent_search.py
@@ -199,17 +199,21 @@
if not agents:
return AgentSearchHits(hits=())
texts: Final = tuple(agent_search_text(agent) for agent in agents)
- cached: Final = self._vectors.get(embedding_model, _NO_VECTORS)
- embedded: Final = await _embed_query_and_agents(embed, query, texts, cached)
+ embedded: Final = await _embed_query_and_agents(
+ embed, query, texts, self._vectors.get(embedding_model, _NO_VECTORS)
+ )
if isinstance(embedded, AgentSearchEmbeddingFailed):
return embedded
if not _same_dimension(embedded.query_vector, embedded.vectors, texts):
return AgentSearchEmbeddingFailed(
reason=f"embedding model {embedding_model} returned vectors of mixed dimensions"
)
- self._vectors = MappingProxyType(
- {**self._vectors, embedding_model: MappingProxyType({**cached, **embedded.vectors})}
+ current: Final = self._vectors.get(embedding_model, _NO_VECTORS)
+ dim: Final = len(embedded.query_vector)
+ merged: Final = MappingProxyType(
+ {text: vector for text, vector in chain(current.items(), embedded.vectors.items()) if len(vector) == dim}
)
+ self._vectors = MappingProxyType({**self._vectors, embedding_model: merged})
ranked: Final = sorted(
(
AgentSearchHit(agent=agent, score=cosine_similarity(embedded.query_vector, embedded.vectors[text]))
diff --git a/tests/test_litellm/proxy/agent_endpoints/test_agent_search.py b/tests/test_litellm/proxy/agent_endpoints/test_agent_search.py
--- a/tests/test_litellm/proxy/agent_endpoints/test_agent_search.py
+++ b/tests/test_litellm/proxy/agent_endpoints/test_agent_search.py
@@ -157,6 +157,19 @@
]
@pytest.mark.asyncio
+ async def test_subset_reembed_after_dimension_change_evicts_stale_vectors(self) -> None:
+ index = AgentSearchIndex()
+ await index.search("language translation", AGENTS, top_k=5, embed=FakeEmbedder(), embedding_model="m")
+ narrow = FixedDimensionEmbedder(2)
+ await index.search("language translation", (TRANSLATOR,), top_k=5, embed=narrow, embedding_model="m")
+ broader = FixedDimensionEmbedder(2)
+ outcome = await index.search("language translation", AGENTS, top_k=5, embed=broader, embedding_model="m")
+ assert isinstance(outcome, AgentSearchHits)
+ assert broader.calls == [
+ ("language translation", agent_search_text(SQL_ANALYST), agent_search_text(TRIP_PLANNER)),
+ ]
+
+ @pytest.mark.asyncio
async def test_mixed_dimensions_in_one_batch_become_embedding_failed(self) -> None:
async def mixed(texts: Sequence[str]) -> Sequence[Vector]:
return ((1.0, 0.0), *((1.0, 0.0, 0.0) for _ in texts[1:]))You can send follow-ups to the cloud agent here.
…ies of another dimension
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 4026aa6. Configure here.
e1cc96e
into
litellm_internal_staging

TLDR
Problem this solves:
How it solves it:
User Flow
Before: a developer whose agents reach other agents through the gateway can only list the registry, so picking the right agent means reading every card by hand
{"name": "agent_search", "arguments": {"query": "translate a pdf document"}}and get back an error naming agent_search an unknown toolAfter: the same developer describes the task in natural language and gets back the best-matching agents their key can reach, ranked
litellm_settings.agent_search_embedding_model: text-embedding-3-small(any embedding model from model_list) and restarts the proxysearch_score{"name": "agent_search", "arguments": {"query": "translate a pdf document"}}and get the ranked agents as JSON: agent_id, name, description, skills, scoreagent_search_not_configurednaming the exact setting to addRelevant issues
Fixes #37829
Linear ticket
Resolves LIT-6309
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Screenshots / Proof of Fix
Shared setup for both legs: two proxy processes share one Postgres DB (
--configbelow, started once each on ports A and B, plus a third process for the last case), so the second-instance case proves the ranking works from a process that never registered the agents. The sameconfig.yamlis used on both sides; the base ignores the unknownlitellm_settingskeyTwo more agents are registered through the API on instance A with the master key, then two virtual keys are minted:
$KEY_SEARCH({"object_permission": {"mcp_tool_search_enabled": true}}) and$KEY_RESTRICTED({"object_permission": {"agents": ["<warehouse-sql-analyst agent_id>"]}})jq -c '.[] | {agent_name, search_score}'printsnullforsearch_scorewherever the response has no such fieldBefore (7083c47)
GET /v1/agents?query= ranks agents
curl -s "http://localhost:$A/v1/agents?query=plan+a+vacation+itinerary&top_k=3" -H "Authorization: Bearer $MASTER" | jq -c '.[] | {agent_name, search_score}'Same query on the second instance
curl -s "http://localhost:$B/v1/agents?query=plan+a+vacation+itinerary&top_k=3" -H "Authorization: Bearer $MASTER" | jq -c '.[] | {agent_name, search_score}'Config-declared agent is searchable
curl -s "http://localhost:$B/v1/agents?query=translate+a+pdf+document&top_k=1" -H "Authorization: Bearer $MASTER" | jq -c '.[] | {agent_name, search_score}'agent_search via POST /mcp-rest/tools/call
curl -s http://localhost:$A/mcp-rest/tools/call -H "Authorization: Bearer $KEY_SEARCH" -H 'Content-Type: application/json' -d '{"name": "agent_search", "arguments": {"query": "check how many units are left in stock", "top_k": 2}}' | jq -c .tools/list and tools/call over POST /mcp/
curl -s http://localhost:$B/mcp/ -H "Authorization: Bearer $KEY_SEARCH" -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' -d '{"jsonrpc": "2.0", "id": 1, "method": "tools/list"}' | grep '^data:' | sed 's/^data: //' | jq -c '[.result.tools[].name]'curl -s http://localhost:$B/mcp/ -H "Authorization: Bearer $KEY_SEARCH" -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' -d '{"jsonrpc": "2.0", "id": 2, "method": "tools/call", "params": {"name": "agent_search", "arguments": {"query": "book a flight and hotel", "top_k": 1}}}' | grep '^data:' | sed 's/^data: //' | jq -c '.result.content[0].text | fromjson? // .'Restricted key only sees its own agents
curl -s "http://localhost:$A/v1/agents?query=plan+a+vacation+itinerary&top_k=3" -H "Authorization: Bearer $KEY_RESTRICTED" | jq -c '.[] | {agent_name, search_score}'No embedding model configured
curl -s -w '\nHTTP %{http_code}\n' "http://localhost:$A/v1/agents?query=plan+a+vacation+itinerary" -H "Authorization: Bearer $MASTER"After (e9cc9c9)
GET /v1/agents?query= ranks agents
curl -s 'http://localhost:$A/v1/agents?query=plan+a+vacation+itinerary&top_k=3' -H 'Authorization: Bearer $MASTER' | jq -c '.[] | {agent_name, search_score}'Same query on the second instance
curl -s 'http://localhost:$B/v1/agents?query=plan+a+vacation+itinerary&top_k=3' -H 'Authorization: Bearer $MASTER' | jq -c '.[] | {agent_name, search_score}'Config-declared agent is searchable
curl -s 'http://localhost:$A/v1/agents?query=translate+a+pdf+document&top_k=1' -H 'Authorization: Bearer $MASTER' | jq -c '.[] | {agent_name, search_score}'agent_search via POST /mcp-rest/tools/call
curl -s http://localhost:$A/mcp-rest/tools/call -H 'Authorization: Bearer $KEY_SEARCH' -H 'Content-Type: application/json' -d '{"name": "agent_search", "arguments": {"query": "check how many units are left in stock", "top_k": 2}}' | jq -c .tools/list and tools/call over POST /mcp/
curl -s http://localhost:$A/mcp/ -H 'Authorization: Bearer $KEY_SEARCH' -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' -d '{"jsonrpc": "2.0", "id": 1, "method": "tools/list"}' | grep '^data:' | sed 's/^data: //' | jq -c '[.result.tools[].name]'curl -s http://localhost:$A/mcp/ -H 'Authorization: Bearer $KEY_SEARCH' -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' -d '{"jsonrpc": "2.0", "id": 2, "method": "tools/call", "params": {"name": "agent_search", "arguments": {"query": "book a flight and hotel", "top_k": 1}}}' | grep '^data:' | sed 's/^data: //' | jq -c '.result.content[0].text | fromjson? // .'curl -s http://localhost:$A/mcp/ -H 'Authorization: Bearer $KEY_SEARCH' -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' -d '{"jsonrpc": "2.0", "id": 3, "method": "tools/call", "params": {"name": "mcp_tool_search", "arguments": {"query": "add numbers"}}}' | grep '^data:' | sed 's/^data: //' | jq -c '{isError: .result.isError, text: .result.content[0].text}'Restricted key only sees its own agents
curl -s 'http://localhost:$A/v1/agents?query=plan+a+vacation+itinerary&top_k=3' -H 'Authorization: Bearer $KEY_RESTRICTED' | jq -c '.[] | {agent_name, search_score}'Embedding spend lands on the calling key
curl -s http://localhost:$A/key/info -H 'Authorization: Bearer $KEY_SEARCH' | jq -c '{key_alias: .info.key_alias, spend: .info.spend}'curl -s 'http://localhost:$A/v1/agents?query=find+the+cheapest+flight+to+paris&top_k=1' -H 'Authorization: Bearer $KEY_SEARCH' | jq -c '[.[] | {agent_name, search_score}]'curl -s http://localhost:$A/key/info -H 'Authorization: Bearer $KEY_SEARCH' | jq -c '{key_alias: .info.key_alias, spend: .info.spend}'curl -s 'http://localhost:$A/spend/logs?api_key=$KEY_SEARCH_HASH' -H 'Authorization: Bearer $MASTER' | jq -c '[.[] | {model, call_type, spend, team_id}]'$KEY_SEARCH_HASHis the.keyfield of the key's own/key/info; every spend log row is an embedding call attributed to itNo embedding model configured
agent_search_embedding_modelline, thencurl -s -w '\nHTTP %{http_code}\n' "http://localhost:$C/v1/agents?query=plan+a+vacation+itinerary" -H "Authorization: Bearer $MASTER"Dependent paths the change touches, same tip
curl -s http://localhost:$A/mcp-rest/tools/list -H 'Authorization: Bearer $KEY_SEARCH' | jq -c '[.tools[] | {name, required: .inputSchema.required}]'requiredis a JSON arraycurl -s http://localhost:$A/mcp-rest/tools/call -H 'Authorization: Bearer $KEY_SEARCH' -H 'Content-Type: application/json' -d '{"name": "mcp_tool_search", "arguments": {"query": "add numbers"}}' | jq -c '{isError, text: .content[0].text}'curl -s http://localhost:$A/mcp-rest/tools/call -H 'Authorization: Bearer $KEY_SEARCH' -H 'Content-Type: application/json' -d '{"name": "agent_search", "arguments": {}}' | jq -c '{isError, text: .content[0].text[0:120]}'curl -s http://localhost:$A/mcp-rest/tools/call -H 'Authorization: Bearer $KEY_SEARCH' -H 'Content-Type: application/json' -d '{"name": "agent_search", "arguments": {"query": "inventory levels", "top_k": "abc"}}' | jq -c '.content[0].text | fromjson | [.[] | {agent_name, score}]'curl -s http://localhost:$A/v1/agents -H 'Authorization: Bearer $MASTER' | jq -c '[.[] | {agent_name, search_score}]'curl -s -o /dev/null -w 'HTTP %{http_code}\n' -X PATCH http://localhost:$A/v1/agents/fc98c971-659a-42dc-902c-b8b8f59a50ef -H 'Authorization: Bearer $MASTER' -H 'Content-Type: application/json' -d '{"agent_name": "trip-planner", "search_score": 0.42}'Before (e9cc9c9) and After (db02cf8): the vector cache survives a change of embedding model
Same DB, keys, and agents as above. Only the config differs: one
agent-embeddermodel group with two deployments of different vector sizes (text-embedding-3-small is 1536-wide, text-embedding-3-large is 3072-wide), which stands in for a router fallback or a config reload onto an embedding model with another vector sizeBefore (e9cc9c9): the second deployment poisons the cache
curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=plan a vacation itinerary' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=translate a pdf document' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=check how many units are left in stock' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=book flights and hotels' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=convert a word file to spanish' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=inventory levels by warehouse' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'After (db02cf8): searches keep answering across both vector sizes
curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=plan a vacation itinerary' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=translate a pdf document' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=check how many units are left in stock' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=book flights and hotels' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=convert a word file to spanish' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=inventory levels by warehouse' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'GET /v1/agents?query= ranks agents
curl -s 'http://localhost:$A/v1/agents?query=plan+a+vacation+itinerary&top_k=3' -H 'Authorization: Bearer $MASTER' | jq -c '.[] | {agent_name, search_score}'Config-declared agent is searchable
curl -s 'http://localhost:$A/v1/agents?query=translate+a+pdf+document&top_k=1' -H 'Authorization: Bearer $MASTER' | jq -c '.[] | {agent_name, search_score}'agent_search via POST /mcp-rest/tools/call
curl -s http://localhost:$A/mcp-rest/tools/call -H 'Authorization: Bearer $KEY_SEARCH' -H 'Content-Type: application/json' -d '{"name": "agent_search", "arguments": {"query": "check how many units are left in stock", "top_k": 2}}' | jq -c .tools/list and tools/call over POST /mcp/
curl -s http://localhost:$A/mcp/ -H 'Authorization: Bearer $KEY_SEARCH' -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' -d '{"jsonrpc": "2.0", "id": 1, "method": "tools/list"}' | grep '^data:' | sed 's/^data: //' | jq -c '[.result.tools[].name]'curl -s http://localhost:$A/mcp/ -H 'Authorization: Bearer $KEY_SEARCH' -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' -d '{"jsonrpc": "2.0", "id": 2, "method": "tools/call", "params": {"name": "agent_search", "arguments": {"query": "book a flight and hotel", "top_k": 1}}}' | grep '^data:' | sed 's/^data: //' | jq -c '.result.content[0].text | fromjson? // .'Restricted key only sees its own agents
curl -s 'http://localhost:$A/v1/agents?query=plan+a+vacation+itinerary&top_k=3' -H 'Authorization: Bearer $KEY_RESTRICTED' | jq -c '.[] | {agent_name, search_score}'Embedding spend lands on the calling key
curl -s http://localhost:$A/key/info -H 'Authorization: Bearer $KEY_SEARCH' | jq -c '{key_alias: .info.key_alias, spend: .info.spend}'curl -s 'http://localhost:$A/v1/agents?query=find+the+cheapest+flight+to+paris&top_k=1' -H 'Authorization: Bearer $KEY_SEARCH' | jq -c '[.[] | {agent_name, search_score}]'curl -s http://localhost:$A/key/info -H 'Authorization: Bearer $KEY_SEARCH' | jq -c '{key_alias: .info.key_alias, spend: .info.spend}'curl -s 'http://localhost:$A/spend/logs?api_key=$KEY_SEARCH_HASH' -H 'Authorization: Bearer $MASTER' | jq -c '[.[] | {model, call_type, spend, team_id}]'$KEY_SEARCH_HASHis the.keyfield of the key's own/key/info; both embedding deployments show up as rows attributed to itBefore (db02cf8) and After (4026aa6): searches running at the same time keep each other's vectors
Same DB and agents as above, back on the single text-embedding-3-small config. Three fresh keys: KEY_TRIP sees trip-planner, KEY_OTHERS sees warehouse-sql-analyst and document-translator, KEY_ALL sees all three. The two restricted keys search at the same time, then KEY_ALL searches twice, and /spend/logs for KEY_ALL shows how many prompt tokens each of its searches embedded (the query alone is 3 tokens)
Before (db02cf8): the first all-agents search pays to embed two agents again
curl -s 'http://localhost:$A/v1/agents?query=book+a+trip&top_k=1' -H "Authorization: Bearer $KEY_TRIP" | jq -c '[.[] | {agent_name, search_score}]' & curl -s 'http://localhost:$A/v1/agents?query=book+a+trip&top_k=1' -H "Authorization: Bearer $KEY_OTHERS" | jq -c '[.[] | {agent_name, search_score}]' & waitcurl -s 'http://localhost:$A/v1/agents?query=book+a+trip&top_k=3' -H "Authorization: Bearer $KEY_ALL" | jq -c '[.[] | {agent_name, search_score}]'curl -s 'http://localhost:$A/v1/agents?query=book+a+trip&top_k=3' -H "Authorization: Bearer $KEY_ALL" | jq -c '[.[] | {agent_name, search_score}]'curl -s 'http://localhost:$A/spend/logs?api_key=$KEY_ALL_HASH' -H 'Authorization: Bearer $MASTER' | jq -c 'sort_by(.startTime) | .[] | {model, call_type, prompt_tokens, spend}'After (4026aa6): the first all-agents search embeds only its query
curl -s 'http://localhost:$A/v1/agents?query=book+a+trip&top_k=1' -H "Authorization: Bearer $KEY_TRIP" | jq -c '[.[] | {agent_name, search_score}]' & curl -s 'http://localhost:$A/v1/agents?query=book+a+trip&top_k=1' -H "Authorization: Bearer $KEY_OTHERS" | jq -c '[.[] | {agent_name, search_score}]' & waitcurl -s 'http://localhost:$A/v1/agents?query=book+a+trip&top_k=3' -H "Authorization: Bearer $KEY_ALL" | jq -c '[.[] | {agent_name, search_score}]'curl -s 'http://localhost:$A/v1/agents?query=book+a+trip&top_k=3' -H "Authorization: Bearer $KEY_ALL" | jq -c '[.[] | {agent_name, search_score}]'curl -s 'http://localhost:$A/spend/logs?api_key=$KEY_ALL_HASH' -H 'Authorization: Bearer $MASTER' | jq -c 'sort_by(.startTime) | .[] | {model, call_type, prompt_tokens, spend}'Dependent paths the change touches, same tip
Back on the two-deployment
agent-embeddergroup from the section above, at 4026aa6: six searches across both vector sizes, agent_search over /mcp-rest, tools/list and tools/call over /mcp/, and a restricted keycurl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=plan a vacation itinerary' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=translate a pdf document' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=check how many units are left in stock' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=book flights and hotels' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=convert a word file to spanish' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=inventory levels by warehouse' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'curl -s http://localhost:$A/mcp-rest/tools/call -H 'Authorization: Bearer $KEY_SEARCH' -H 'Content-Type: application/json' -d '{"name": "agent_search", "arguments": {"query": "check how many units are left in stock", "top_k": 2}}' | jq -c .curl -s http://localhost:$A/mcp/ -H 'Authorization: Bearer $KEY_SEARCH' -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' -d '{"jsonrpc": "2.0", "id": 1, "method": "tools/list"}' | grep '^data:' | sed 's/^data: //' | jq -c '[.result.tools[].name]'curl -s http://localhost:$A/mcp/ -H 'Authorization: Bearer $KEY_SEARCH' -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' -d '{"jsonrpc": "2.0", "id": 2, "method": "tools/call", "params": {"name": "agent_search", "arguments": {"query": "book a flight and hotel", "top_k": 1}}}' | grep '^data:' | sed 's/^data: //' | jq -c '.result.content[0].text | fromjson? // .'curl -s 'http://localhost:$A/v1/agents?query=plan+a+vacation+itinerary&top_k=3' -H 'Authorization: Bearer $KEY_RESTRICTED' | jq -c '.[] | {agent_name, search_score}'Type
🆕 New Feature
Caveats (if any)
Low
Final Attestation
The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR
4026aa6 passes /live-pr-risk
Note
Medium Risk
Search triggers billable embedding calls and merges a global in-process vector cache, but behavior is scoped to agent discovery, reuses existing key RBAC, and fails explicitly when misconfigured.
Overview
Adds semantic ranking over the A2A agent registry so callers can describe a task in natural language instead of scanning every agent card.
REST:
GET /v1/agentsnow accepts optionalqueryandtop_k. Whenqueryis set, accessible agents are embedded and ranked by cosine similarity; each row gets asearch_score. Missinglitellm_settings.agent_search_embedding_modelreturns 400; embedding failures return 503. Listing withoutqueryis unchanged (scores stay null).Config: New
agent_search_embedding_modelselects an embedding model frommodel_list. Embedding spend is attributed to the calling API key via router metadata.Shared core: New
agent_searchmodule implements text extraction from agent cards, router-backed embeddings, a per-model vector cache (handles mixed vector dimensions and concurrent searches), andsearch_agents()used by both REST and MCP.MCP: Virtual tool
agent_searchjoinsmcp_tool_search/mcp_tool_callbehindVIRTUAL_TOOL_NAMES, still gated onmcp_tool_search_enabled. REST and protocol MCP paths dispatch to the same handler and return ranked JSON (or tool errors when not configured).RBAC: Agent listing and search use shared
accessible_agents()so results only include agents the key (and team grants) can reach.OpenAPI / dashboard
schema.d.tsare updated for the new query params,search_score, and virtual tool schema typing.Reviewed by Cursor Bugbot for commit 4026aa6. Bugbot is set up for automated code reviews on this repo. Configure here.