Skip to content

feat(a2a): semantic search over the agent registry via GET /v1/agents?query and an agent_search MCP tool - #38609

Merged
mateo-berri merged 6 commits into
litellm_internal_stagingfrom
litellm_a2a_agent_semantic_search
Aug 28, 2026
Merged

mateo-berri merged 6 commits into
litellm_internal_stagingfrom
litellm_a2a_agent_semantic_search

Conversation

@mateo-berri

@mateo-berri mateo-berri commented Aug 28, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Agent registries grow past what an agent can scan by name
  • GET /v1/agents only lists; there is no way to search it
  • MCP clients get mcp_tool_search for tools but nothing for agents

How it solves it:

  • GET /v1/agents accepts query and top_k, ranks by semantic similarity
  • New agent_search MCP virtual tool for keys with tool search enabled
  • One config key: litellm_settings.agent_search_embedding_model picks the embedding model
  • Results only cover agents the calling key can access
  • Ranked responses carry a search_score field per agent
  • Embedding spend is tracked against the calling key, so /key/info and /spend/logs see it
  • Agent vectors are cached per embedding model and re-embedded when the vector size changes, so a fallback or a config change to another embedding model never breaks search
  • Fresh agent vectors merge into the live cache, so searches running at the same time keep each other's work and the next search never pays to embed an agent twice

User Flow

Before: a developer whose agents reach other agents through the gateway can only list the registry, so picking the right agent means reading every card by hand

  1. They send GET https://litellm-domain/v1/agents?query=translate+a+pdf+document with their key and get 200 with the whole registry, unranked; the query parameter is silently ignored
  2. They send POST https://litellm-domain/mcp-rest/tools/call with {"name": "agent_search", "arguments": {"query": "translate a pdf document"}} and get back an error naming agent_search an unknown tool
  3. Their MCP client lists tools over POST https://litellm-domain/mcp/ and sees only mcp_tool_search and mcp_tool_call, neither of which covers agents

After: the same developer describes the task in natural language and gets back the best-matching agents their key can reach, ranked

  1. The proxy admin sets litellm_settings.agent_search_embedding_model: text-embedding-3-small (any embedding model from model_list) and restarts the proxy
  2. They send GET https://litellm-domain/v1/agents?query=translate+a+pdf+document&top_k=3 and get 200 with up to 3 agents ranked by semantic similarity, each carrying a search_score
  3. They send POST https://litellm-domain/mcp-rest/tools/call with {"name": "agent_search", "arguments": {"query": "translate a pdf document"}} and get the ranked agents as JSON: agent_id, name, description, skills, score
  4. Their MCP client lists tools over POST https://litellm-domain/mcp/, sees agent_search alongside the two existing tools, and tools/call on it returns the same ranked JSON
  5. A key restricted to specific agents only ever gets its own agents back, whatever the query says
  6. If the admin never set the config key, GET /v1/agents?query=... answers 400 agent_search_not_configured naming the exact setting to add
  7. The embedding calls those searches make show up under their key: GET https://litellm-domain/key/info spend grows and GET https://litellm-domain/spend/logs?api_key= lists them
  8. If the admin later points agent_search_embedding_model at a model with a different vector size, or the router falls back to one, the next GET https://litellm-domain/v1/agents?query=... still answers 200 with a ranked list
  9. Two keys that see different agents can search at the same time; the next search from a key that sees every agent embeds only its query, so GET https://litellm-domain/spend/logs?api_key= shows 3 prompt tokens for it, not the whole registry

Relevant issues

Fixes #37829

Linear ticket

Resolves LIT-6309

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Screenshots / Proof of Fix

Shared setup for both legs: two proxy processes share one Postgres DB (--config below, started once each on ports A and B, plus a third process for the last case), so the second-instance case proves the ranking works from a process that never registered the agents. The same config.yaml is used on both sides; the base ignores the unknown litellm_settings key

model_list:
  - model_name: text-embedding-3-small
    litellm_params:
      model: openai/text-embedding-3-small
      api_key: os.environ/OPENAI_API_KEY
litellm_settings:
  agent_search_embedding_model: text-embedding-3-small
general_settings:
  store_model_in_db: true
agents:
  - agent_name: document-translator
    agent_card_params:
      name: Document Translator
      description: Converts PDF, Word and text files from one language into another while preserving layout
      url: http://localhost:10001
      protocolVersion: "0.3"
      version: "1.0.0"
      capabilities: {}
      defaultInputModes: ["text"]
      defaultOutputModes: ["text"]
      skills:
        - id: translate-file
          name: Translate a file
          description: Take an uploaded document and produce the same document in the target language
          tags: ["localization", "documents"]

Two more agents are registered through the API on instance A with the master key, then two virtual keys are minted: $KEY_SEARCH ({"object_permission": {"mcp_tool_search_enabled": true}}) and $KEY_RESTRICTED ({"object_permission": {"agents": ["<warehouse-sql-analyst agent_id>"]}})

curl -s http://localhost:$A/v1/agents -H "Authorization: Bearer $MASTER" -H 'Content-Type: application/json' -d '{"agent_name": "trip-planner", "agent_card_params": {"protocolVersion": "0.3", "name": "Trip Planner", "description": "Books flights and hotels for a business trip given dates, origin and destination", "url": "http://localhost:10003", "version": "1.0.0", "capabilities": {}, "defaultInputModes": ["text"], "defaultOutputModes": ["text"], "skills": [{"id": "book-travel", "name": "Book travel", "description": "Find and book the flights and hotel for a business trip", "tags": ["travel", "booking"]}]}}'
curl -s http://localhost:$A/v1/agents -H "Authorization: Bearer $MASTER" -H 'Content-Type: application/json' -d '{"agent_name": "warehouse-sql-analyst", "agent_card_params": {"protocolVersion": "0.3", "name": "Warehouse SQL Analyst", "description": "Answers questions about stock by running SQL queries against the inventory database", "url": "http://localhost:10002", "version": "1.0.0", "capabilities": {}, "defaultInputModes": ["text"], "defaultOutputModes": ["text"], "skills": [{"id": "query-stock-levels", "name": "Query stock levels", "description": "Run a SQL query against the inventory database and summarize the stock levels it returns", "tags": ["sql", "analytics"]}]}}'

jq -c '.[] | {agent_name, search_score}' prints null for search_score wherever the response has no such field

Before (7083c47)

GET /v1/agents?query= ranks agents

  1. curl -s "http://localhost:$A/v1/agents?query=plan+a+vacation+itinerary&top_k=3" -H "Authorization: Bearer $MASTER" | jq -c '.[] | {agent_name, search_score}'
  2. The query and top_k are ignored: the full registry comes back in insertion order with no score
    {"agent_name":"trip-planner","search_score":null}
    {"agent_name":"warehouse-sql-analyst","search_score":null}
    {"agent_name":"document-translator","search_score":null}
    

Same query on the second instance

  1. curl -s "http://localhost:$B/v1/agents?query=plan+a+vacation+itinerary&top_k=3" -H "Authorization: Bearer $MASTER" | jq -c '.[] | {agent_name, search_score}'
  2. Same unranked list
    {"agent_name":"trip-planner","search_score":null}
    {"agent_name":"warehouse-sql-analyst","search_score":null}
    {"agent_name":"document-translator","search_score":null}
    

Config-declared agent is searchable

  1. curl -s "http://localhost:$B/v1/agents?query=translate+a+pdf+document&top_k=1" -H "Authorization: Bearer $MASTER" | jq -c '.[] | {agent_name, search_score}'
  2. top_k=1 is ignored too; all three agents come back unranked
    {"agent_name":"trip-planner","search_score":null}
    {"agent_name":"warehouse-sql-analyst","search_score":null}
    {"agent_name":"document-translator","search_score":null}
    

agent_search via POST /mcp-rest/tools/call

  1. curl -s http://localhost:$A/mcp-rest/tools/call -H "Authorization: Bearer $KEY_SEARCH" -H 'Content-Type: application/json' -d '{"name": "agent_search", "arguments": {"query": "check how many units are left in stock", "top_k": 2}}' | jq -c .
  2. agent_search is not a virtual tool, so the REST route treats it as a regular MCP tool and rejects the call
    {"detail":{"error":"missing_parameter","message":"server_id is required in request body"}}
    

tools/list and tools/call over POST /mcp/

  1. curl -s http://localhost:$B/mcp/ -H "Authorization: Bearer $KEY_SEARCH" -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' -d '{"jsonrpc": "2.0", "id": 1, "method": "tools/list"}' | grep '^data:' | sed 's/^data: //' | jq -c '[.result.tools[].name]'
  2. Only the two existing virtual tools are listed
    ["mcp_tool_search","mcp_tool_call"]
    
  3. curl -s http://localhost:$B/mcp/ -H "Authorization: Bearer $KEY_SEARCH" -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' -d '{"jsonrpc": "2.0", "id": 2, "method": "tools/call", "params": {"name": "agent_search", "arguments": {"query": "book a flight and hotel", "top_k": 1}}}' | grep '^data:' | sed 's/^data: //' | jq -c '.result.content[0].text | fromjson? // .'
  4. The call is refused
    "Error: User not allowed to call this tool."
    

Restricted key only sees its own agents

  1. curl -s "http://localhost:$A/v1/agents?query=plan+a+vacation+itinerary&top_k=3" -H "Authorization: Bearer $KEY_RESTRICTED" | jq -c '.[] | {agent_name, search_score}'
  2. The key's one agent comes back, unscored
    {"agent_name":"warehouse-sql-analyst","search_score":null}
    

No embedding model configured

  1. curl -s -w '\nHTTP %{http_code}\n' "http://localhost:$A/v1/agents?query=plan+a+vacation+itinerary" -H "Authorization: Bearer $MASTER"
  2. Nothing tells the caller that search is unavailable: 200 with the full registry (payload elided, it is the same three agents)
    [{"agent_id":"e6210b81-5505-4595-920b-53288f704809","agent_name":"trip-planner", ...}, {"agent_id":"ad1c73ee-deff-4d1c-85f7-a42fbc1ec255","agent_name":"warehouse-sql-analyst", ...}, {"agent_id":"a9cb3976...","agent_name":"document-translator", ...}]
    HTTP 200
    

After (e9cc9c9)

GET /v1/agents?query= ranks agents

  1. curl -s 'http://localhost:$A/v1/agents?query=plan+a+vacation+itinerary&top_k=3' -H 'Authorization: Bearer $MASTER' | jq -c '.[] | {agent_name, search_score}'
  2. trip-planner ranks first and every agent carries its cosine score
    {"agent_name":"trip-planner","search_score":0.46068230836959995}
    {"agent_name":"document-translator","search_score":0.11200725760167635}
    {"agent_name":"warehouse-sql-analyst","search_score":0.09470067778779014}
    

Same query on the second instance

  1. curl -s 'http://localhost:$B/v1/agents?query=plan+a+vacation+itinerary&top_k=3' -H 'Authorization: Bearer $MASTER' | jq -c '.[] | {agent_name, search_score}'
  2. Instance B never registered the two API agents; it embeds them from the shared DB and ranks identically
    {"agent_name":"trip-planner","search_score":0.4607150537704817}
    {"agent_name":"document-translator","search_score":0.11204450076080524}
    {"agent_name":"warehouse-sql-analyst","search_score":0.094646121092808}
    

Config-declared agent is searchable

  1. curl -s 'http://localhost:$A/v1/agents?query=translate+a+pdf+document&top_k=1' -H 'Authorization: Bearer $MASTER' | jq -c '.[] | {agent_name, search_score}'
  2. The agent declared in config.yaml is ranked like the DB ones, and top_k=1 trims the list
    {"agent_name":"document-translator","search_score":0.6732450375047712}
    

agent_search via POST /mcp-rest/tools/call

  1. curl -s http://localhost:$A/mcp-rest/tools/call -H 'Authorization: Bearer $KEY_SEARCH' -H 'Content-Type: application/json' -d '{"name": "agent_search", "arguments": {"query": "check how many units are left in stock", "top_k": 2}}' | jq -c .
  2. The ranked agents come back as JSON text with agent_id, name, description, skills, and score
    {"_meta":null,"content":[{"type":"text","text":"[{\"agent_id\": \"b21b8787-8b9e-4c5f-a45c-3f5e4061d70e\", \"agent_name\": \"warehouse-sql-analyst\", \"description\": \"Answers questions about stock by running SQL queries against the inventory database\", \"skills\": [{\"name\": \"Query stock levels\", \"description\": \"Run a SQL query against the inventory database and summarize the stock levels it returns\", \"tags\": [\"sql\", \"analytics\"]}], \"score\": 0.3791485568539873}, {\"agent_id\": \"fc98c971-659a-42dc-902c-b8b8f59a50ef\", \"agent_name\": \"trip-planner\", \"description\": \"Books flights and hotels for a business trip given dates, origin and destination\", \"skills\": [{\"name\": \"Book travel\", \"description\": \"Find and book the flights and hotel for a business trip\", \"tags\": [\"travel\", \"booking\"]}], \"score\": 0.12168402727229097}]","annotations":null,"_meta":null}],"structuredContent":null,"isError":false}
    

tools/list and tools/call over POST /mcp/

  1. curl -s http://localhost:$A/mcp/ -H 'Authorization: Bearer $KEY_SEARCH' -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' -d '{"jsonrpc": "2.0", "id": 1, "method": "tools/list"}' | grep '^data:' | sed 's/^data: //' | jq -c '[.result.tools[].name]'
  2. agent_search is listed next to the two existing virtual tools
    ["mcp_tool_search","mcp_tool_call","agent_search"]
    
  3. curl -s http://localhost:$A/mcp/ -H 'Authorization: Bearer $KEY_SEARCH' -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' -d '{"jsonrpc": "2.0", "id": 2, "method": "tools/call", "params": {"name": "agent_search", "arguments": {"query": "book a flight and hotel", "top_k": 1}}}' | grep '^data:' | sed 's/^data: //' | jq -c '.result.content[0].text | fromjson? // .'
  4. The MCP server's jsonschema validation passes and the best agent comes back
    [{"agent_id":"fc98c971-659a-42dc-902c-b8b8f59a50ef","agent_name":"trip-planner","description":"Books flights and hotels for a business trip given dates, origin and destination","skills":[{"name":"Book travel","description":"Find and book the flights and hotel for a business trip","tags":["travel","booking"]}],"score":0.6081425046921755}]
    
  5. curl -s http://localhost:$A/mcp/ -H 'Authorization: Bearer $KEY_SEARCH' -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' -d '{"jsonrpc": "2.0", "id": 3, "method": "tools/call", "params": {"name": "mcp_tool_search", "arguments": {"query": "add numbers"}}}' | grep '^data:' | sed 's/^data: //' | jq -c '{isError: .result.isError, text: .result.content[0].text}'
  6. The pre-existing mcp_tool_search still validates and answers (no MCP servers are configured, so the match list is empty)
    {"isError":false,"text":"[]"}
    

Restricted key only sees its own agents

  1. curl -s 'http://localhost:$A/v1/agents?query=plan+a+vacation+itinerary&top_k=3' -H 'Authorization: Bearer $KEY_RESTRICTED' | jq -c '.[] | {agent_name, search_score}'
  2. trip-planner is the best match but this key cannot see it; only its own agent is ranked
    {"agent_name":"warehouse-sql-analyst","search_score":0.09472462377311118}
    

Embedding spend lands on the calling key

  1. curl -s http://localhost:$A/key/info -H 'Authorization: Bearer $KEY_SEARCH' | jq -c '{key_alias: .info.key_alias, spend: .info.spend}'
  2. The search key starts with no spend
    {"key_alias":null,"spend":0.0}
    
  3. curl -s 'http://localhost:$A/v1/agents?query=find+the+cheapest+flight+to+paris&top_k=1' -H 'Authorization: Bearer $KEY_SEARCH' | jq -c '[.[] | {agent_name, search_score}]'
  4. One ranked search with that key
    [{"agent_name":"trip-planner","search_score":0.30937728168379236}]
    
  5. curl -s http://localhost:$A/key/info -H 'Authorization: Bearer $KEY_SEARCH' | jq -c '{key_alias: .info.key_alias, spend: .info.spend}'
  6. 20 seconds later the key's spend includes the embedding calls the searches on this process made
    {"key_alias":null,"spend":3.8E-7}
    
  7. curl -s 'http://localhost:$A/spend/logs?api_key=$KEY_SEARCH_HASH' -H 'Authorization: Bearer $MASTER' | jq -c '[.[] | {model, call_type, spend, team_id}]'
  8. $KEY_SEARCH_HASH is the .key field of the key's own /key/info; every spend log row is an embedding call attributed to it
    [{"model":"openai/text-embedding-3-small","call_type":"aembedding","spend":1.2E-7,"team_id":""},{"model":"openai/text-embedding-3-small","call_type":"aembedding","spend":1E-7,"team_id":""},{"model":"openai/text-embedding-3-small","call_type":"aembedding","spend":1.6E-7,"team_id":""}]
    

No embedding model configured

  1. Start a third process from the same config minus the agent_search_embedding_model line, then curl -s -w '\nHTTP %{http_code}\n' "http://localhost:$C/v1/agents?query=plan+a+vacation+itinerary" -H "Authorization: Bearer $MASTER"
  2. The caller is told exactly which setting is missing
    {"detail":{"error":"agent_search_not_configured","message":"agent search needs litellm_settings.agent_search_embedding_model set to an embedding model from model_list"}}
    HTTP 400
    

Dependent paths the change touches, same tip

  1. curl -s http://localhost:$A/mcp-rest/tools/list -H 'Authorization: Bearer $KEY_SEARCH' | jq -c '[.tools[] | {name, required: .inputSchema.required}]'
  2. GET /mcp-rest/tools/list serializes the virtual tool tuple and every required is a JSON array
    [{"name":"mcp_tool_search","required":["query"]},{"name":"mcp_tool_call","required":["tool_name"]},{"name":"agent_search","required":["query"]}]
    
  3. curl -s http://localhost:$A/mcp-rest/tools/call -H 'Authorization: Bearer $KEY_SEARCH' -H 'Content-Type: application/json' -d '{"name": "mcp_tool_search", "arguments": {"query": "add numbers"}}' | jq -c '{isError, text: .content[0].text}'
  4. mcp_tool_search still dispatches over /mcp-rest
    {"isError":false,"text":"[]"}
    
  5. curl -s http://localhost:$A/mcp-rest/tools/call -H 'Authorization: Bearer $KEY_SEARCH' -H 'Content-Type: application/json' -d '{"name": "agent_search", "arguments": {}}' | jq -c '{isError, text: .content[0].text[0:120]}'
  6. Over /mcp-rest nothing validates arguments, so a missing query reaches the embedding model and comes back as a tool error, never a 500
    {"isError":true,"text":"embedding the search query failed: litellm.BadRequestError: OpenAIException - Error code: 400 - {'error': {'message': \"I"}
    
  7. curl -s http://localhost:$A/mcp-rest/tools/call -H 'Authorization: Bearer $KEY_SEARCH' -H 'Content-Type: application/json' -d '{"name": "agent_search", "arguments": {"query": "inventory levels", "top_k": "abc"}}' | jq -c '.content[0].text | fromjson | [.[] | {agent_name, score}]'
  8. A non-numeric top_k falls back to the default of 5
    [{"agent_name":"warehouse-sql-analyst","score":0.4423848043943892},{"agent_name":"trip-planner","score":0.1824006902772436},{"agent_name":"document-translator","score":0.11647182300071149}]
    
  9. curl -s http://localhost:$A/v1/agents -H 'Authorization: Bearer $MASTER' | jq -c '[.[] | {agent_name, search_score}]'
  10. Without query the list is unranked and search_score is null
[{"agent_name":"warehouse-sql-analyst","search_score":null},{"agent_name":"trip-planner","search_score":null},{"agent_name":"document-translator","search_score":null}]
  1. curl -s -o /dev/null -w 'HTTP %{http_code}\n' -X PATCH http://localhost:$A/v1/agents/fc98c971-659a-42dc-902c-b8b8f59a50ef -H 'Authorization: Bearer $MASTER' -H 'Content-Type: application/json' -d '{"agent_name": "trip-planner", "search_score": 0.42}'
  2. PATCH /v1/agents/{id} with a body carrying search_score (the dashboard round-trips whole agent objects) still writes; the field never reaches Prisma
HTTP 200

Before (e9cc9c9) and After (db02cf8): the vector cache survives a change of embedding model

Same DB, keys, and agents as above. Only the config differs: one agent-embedder model group with two deployments of different vector sizes (text-embedding-3-small is 1536-wide, text-embedding-3-large is 3072-wide), which stands in for a router fallback or a config reload onto an embedding model with another vector size

model_list:
  - model_name: agent-embedder
    litellm_params:
      model: openai/text-embedding-3-small
      api_key: os.environ/OPENAI_API_KEY
  - model_name: agent-embedder
    litellm_params:
      model: openai/text-embedding-3-large
      api_key: os.environ/OPENAI_API_KEY
litellm_settings:
  agent_search_embedding_model: agent-embedder

Before (e9cc9c9): the second deployment poisons the cache

  1. curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=plan a vacation itinerary' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'
  2. The first search embeds the agents through whichever deployment the router picks and caches those vectors
    [{"agent_name":"trip-planner","search_score":0.4297847119797013}]
    
  3. curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=translate a pdf document' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'
  4. A search whose query lands on the other deployment compares a 3072-wide query against 1536-wide cached vectors and dies with a 500
    {"detail":{"error":"Internal server error: zip() argument 2 is longer than argument 1"}}
    
  5. curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=check how many units are left in stock' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'
  6. Same again; nothing evicts the poisoned cache
    {"detail":{"error":"Internal server error: zip() argument 2 is longer than argument 1"}}
    
  7. curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=book flights and hotels' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'
  8. A search that happens to land on the cached deployment still works, so the failure looks random to the caller
    [{"agent_name":"trip-planner","search_score":0.5967593470541304}]
    
  9. curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=convert a word file to spanish' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'
  10. Works when the router picks the cached deployment
[{"agent_name":"document-translator","search_score":0.5123179525580667}]
  1. curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=inventory levels by warehouse' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'
  2. Works for the same reason; the 500s keep coming for as long as the worker lives
[{"agent_name":"warehouse-sql-analyst","search_score":0.49354130916348515}]

After (db02cf8): searches keep answering across both vector sizes

  1. curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=plan a vacation itinerary' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'
  2. First search fills the cache for this embedding model
    [{"agent_name":"trip-planner","search_score":0.46070815142107585}]
    
  3. curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=translate a pdf document' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'
  4. A query embedded at the other vector size re-embeds the query and the agents together in one call, then ranks
    [{"agent_name":"document-translator","search_score":0.6732450375047712}]
    
  5. curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=check how many units are left in stock' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'
  6. Ranks
    [{"agent_name":"warehouse-sql-analyst","search_score":0.3771639080808998}]
    
  7. curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=book flights and hotels' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'
  8. Ranks
    [{"agent_name":"trip-planner","search_score":0.5967593470541304}]
    
  9. curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=convert a word file to spanish' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'
  10. Ranks
[{"agent_name":"document-translator","search_score":0.5382340353586434}]
  1. curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=inventory levels by warehouse' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'
  2. Six searches over a group that flips between 1536 and 3072 dimensions, six ranked answers, no 500 and no 503
[{"agent_name":"warehouse-sql-analyst","search_score":0.5334472063220455}]

GET /v1/agents?query= ranks agents

  1. curl -s 'http://localhost:$A/v1/agents?query=plan+a+vacation+itinerary&top_k=3' -H 'Authorization: Bearer $MASTER' | jq -c '.[] | {agent_name, search_score}'
  2. trip-planner ranks first and every agent carries its cosine score
    {"agent_name":"trip-planner","search_score":0.46070815142107585}
    {"agent_name":"document-translator","search_score":0.11203058136724742}
    {"agent_name":"warehouse-sql-analyst","search_score":0.09471439825275152}
    

Config-declared agent is searchable

  1. curl -s 'http://localhost:$A/v1/agents?query=translate+a+pdf+document&top_k=1' -H 'Authorization: Bearer $MASTER' | jq -c '.[] | {agent_name, search_score}'
  2. The yaml-declared translator still ranks first for its own task
    {"agent_name":"document-translator","search_score":0.6732509313613476}
    

agent_search via POST /mcp-rest/tools/call

  1. curl -s http://localhost:$A/mcp-rest/tools/call -H 'Authorization: Bearer $KEY_SEARCH' -H 'Content-Type: application/json' -d '{"name": "agent_search", "arguments": {"query": "check how many units are left in stock", "top_k": 2}}' | jq -c .
  2. The MCP virtual tool ranks the same registry through the same cache
    {"_meta":null,"content":[{"type":"text","text":"[{\"agent_id\": \"b21b8787-8b9e-4c5f-a45c-3f5e4061d70e\", \"agent_name\": \"warehouse-sql-analyst\", \"description\": \"Answers questions about stock by running SQL queries against the inventory database\", \"skills\": [{\"name\": \"Query stock levels\", \"description\": \"Run a SQL query against the inventory database and summarize the stock levels it returns\", \"tags\": [\"sql\", \"analytics\"]}], \"score\": 0.377168456111392}, {\"agent_id\": \"fc98c971-659a-42dc-902c-b8b8f59a50ef\", \"agent_name\": \"trip-planner\", \"description\": \"Books flights and hotels for a business trip given dates, origin and destination\", \"skills\": [{\"name\": \"Book travel\", \"description\": \"Find and book the flights and hotel for a business trip\", \"tags\": [\"travel\", \"booking\"]}], \"score\": 0.12617954845303345}]","annotations":null,"_meta":null}],"structuredContent":null,"isError":false}
    

tools/list and tools/call over POST /mcp/

  1. curl -s http://localhost:$A/mcp/ -H 'Authorization: Bearer $KEY_SEARCH' -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' -d '{"jsonrpc": "2.0", "id": 1, "method": "tools/list"}' | grep '^data:' | sed 's/^data: //' | jq -c '[.result.tools[].name]'
  2. agent_search is listed next to the two existing virtual tools
    ["mcp_tool_search","mcp_tool_call","agent_search"]
    
  3. curl -s http://localhost:$A/mcp/ -H 'Authorization: Bearer $KEY_SEARCH' -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' -d '{"jsonrpc": "2.0", "id": 2, "method": "tools/call", "params": {"name": "agent_search", "arguments": {"query": "book a flight and hotel", "top_k": 1}}}' | grep '^data:' | sed 's/^data: //' | jq -c '.result.content[0].text | fromjson? // .'
  4. tools/call on it returns the ranked JSON
    [{"agent_id":"fc98c971-659a-42dc-902c-b8b8f59a50ef","agent_name":"trip-planner","description":"Books flights and hotels for a business trip given dates, origin and destination","skills":[{"name":"Book travel","description":"Find and book the flights and hotel for a business trip","tags":["travel","booking"]}],"score":0.5697225709111288}]
    

Restricted key only sees its own agents

  1. curl -s 'http://localhost:$A/v1/agents?query=plan+a+vacation+itinerary&top_k=3' -H 'Authorization: Bearer $KEY_RESTRICTED' | jq -c '.[] | {agent_name, search_score}'
  2. A key limited to the warehouse agent gets only that agent back, whatever the query
    {"agent_name":"warehouse-sql-analyst","search_score":0.05868779048073312}
    

Embedding spend lands on the calling key

  1. curl -s http://localhost:$A/key/info -H 'Authorization: Bearer $KEY_SEARCH' | jq -c '{key_alias: .info.key_alias, spend: .info.spend}'
  2. The search key's spend before this search
    {"key_alias":null,"spend":0.00002451}
    
  3. curl -s 'http://localhost:$A/v1/agents?query=find+the+cheapest+flight+to+paris&top_k=1' -H 'Authorization: Bearer $KEY_SEARCH' | jq -c '[.[] | {agent_name, search_score}]'
  4. One ranked search with that key
    [{"agent_name":"trip-planner","search_score":0.3307554907253355}]
    
  5. curl -s http://localhost:$A/key/info -H 'Authorization: Bearer $KEY_SEARCH' | jq -c '{key_alias: .info.key_alias, spend: .info.spend}'
  6. 20 seconds later the key's spend has grown by the embedding calls, including the re-embed after a vector-size flip
    {"key_alias":null,"spend":0.0000575}
    
  7. curl -s 'http://localhost:$A/spend/logs?api_key=$KEY_SEARCH_HASH' -H 'Authorization: Bearer $MASTER' | jq -c '[.[] | {model, call_type, spend, team_id}]'
  8. $KEY_SEARCH_HASH is the .key field of the key's own /key/info; both embedding deployments show up as rows attributed to it
    [{"model":"openai/text-embedding-3-large","call_type":"aembedding","spend":7.799999999999999E-7,"team_id":""},{"model":"openai/text-embedding-3-large","call_type":"aembedding","spend":0.00001534,"team_id":""},{"model":"openai/text-embedding-3-small","call_type":"aembedding","spend":1E-7,"team_id":""},{"model":"openai/text-embedding-3-large","call_type":"aembedding","spend":0.00001573,"team_id":""},{"model":"openai/text-embedding-3-large","call_type":"aembedding","spend":0.00000104,"team_id":""},{"model":"openai/text-embedding-3-small","call_type":"aembedding","spend":0.00000238,"team_id":""},{"model":"openai/text-embedding-3-large","call_type":"aembedding","spend":7.799999999999999E-7,"team_id":""},{"model":"openai/text-embedding-3-small","call_type":"aembedding","spend":1E-7,"team_id":""},{"model":"openai/text-embedding-3-small","call_type":"aembedding","spend":0.00000242,"team_id":""},{"model":"openai/text-embedding-3-large","call_type":"aembedding","spend":0.00000104,"team_id":""},{"model":"openai/text-embedding-3-small","call_type":"aembedding","spend":8E-7,"team_id":""},{"model":"openai/text-embedding-3-large","call_type":"aembedding","spend":7.799999999999999E-7,"team_id":""},{"model":"openai/text-embedding-3-large","call_type":"aembedding","spend":0.00001469,"team_id":""},{"model":"openai/text-embedding-3-small","call_type":"aembedding","spend":1E-7,"team_id":""},{"model":"openai/text-embedding-3-large","call_type":"aembedding","spend":0.00000104,"team_id":""},{"model":"openai/text-embedding-3-small","call_type":"aembedding","spend":0.0000023,"team_id":""},{"model":"openai/text-embedding-3-small","call_type":"aembedding","spend":1.2E-7,"team_id":""},{"model":"openai/text-embedding-3-small","call_type":"aembedding","spend":1E-7,"team_id":""},{"model":"openai/text-embedding-3-small","call_type":"aembedding","spend":1.6E-7,"team_id":""}]
    

Before (db02cf8) and After (4026aa6): searches running at the same time keep each other's vectors

Same DB and agents as above, back on the single text-embedding-3-small config. Three fresh keys: KEY_TRIP sees trip-planner, KEY_OTHERS sees warehouse-sql-analyst and document-translator, KEY_ALL sees all three. The two restricted keys search at the same time, then KEY_ALL searches twice, and /spend/logs for KEY_ALL shows how many prompt tokens each of its searches embedded (the query alone is 3 tokens)

Before (db02cf8): the first all-agents search pays to embed two agents again

  1. curl -s 'http://localhost:$A/v1/agents?query=book+a+trip&top_k=1' -H "Authorization: Bearer $KEY_TRIP" | jq -c '[.[] | {agent_name, search_score}]' & curl -s 'http://localhost:$A/v1/agents?query=book+a+trip&top_k=1' -H "Authorization: Bearer $KEY_OTHERS" | jq -c '[.[] | {agent_name, search_score}]' & wait
    [{"agent_name":"document-translator","search_score":0.15166149306363558}]
    [{"agent_name":"trip-planner","search_score":0.6536557773948982}]
    
  2. curl -s 'http://localhost:$A/v1/agents?query=book+a+trip&top_k=3' -H "Authorization: Bearer $KEY_ALL" | jq -c '[.[] | {agent_name, search_score}]'
    [{"agent_name":"trip-planner","search_score":0.653619854480488},{"agent_name":"document-translator","search_score":0.15166149306363558},{"agent_name":"warehouse-sql-analyst","search_score":0.1008761074424888}]
    
  3. curl -s 'http://localhost:$A/v1/agents?query=book+a+trip&top_k=3' -H "Authorization: Bearer $KEY_ALL" | jq -c '[.[] | {agent_name, search_score}]'
    [{"agent_name":"trip-planner","search_score":0.6537126920899458},{"agent_name":"document-translator","search_score":0.15167465008416967},{"agent_name":"warehouse-sql-analyst","search_score":0.10085297030343146}]
    
  4. curl -s 'http://localhost:$A/spend/logs?api_key=$KEY_ALL_HASH' -H 'Authorization: Bearer $MASTER' | jq -c 'sort_by(.startTime) | .[] | {model, call_type, prompt_tokens, spend}'
    {"model":"openai/text-embedding-3-small","call_type":"aembedding","prompt_tokens":82,"spend":0.00000164}
    {"model":"openai/text-embedding-3-small","call_type":"aembedding","prompt_tokens":3,"spend":6.000000000000001E-8}
    

After (4026aa6): the first all-agents search embeds only its query

  1. curl -s 'http://localhost:$A/v1/agents?query=book+a+trip&top_k=1' -H "Authorization: Bearer $KEY_TRIP" | jq -c '[.[] | {agent_name, search_score}]' & curl -s 'http://localhost:$A/v1/agents?query=book+a+trip&top_k=1' -H "Authorization: Bearer $KEY_OTHERS" | jq -c '[.[] | {agent_name, search_score}]' & wait
    [{"agent_name":"document-translator","search_score":0.1516644645463074}]
    [{"agent_name":"trip-planner","search_score":0.6536557773948982}]
    
  2. curl -s 'http://localhost:$A/v1/agents?query=book+a+trip&top_k=3' -H "Authorization: Bearer $KEY_ALL" | jq -c '[.[] | {agent_name, search_score}]'
    [{"agent_name":"trip-planner","search_score":0.653648258815919},{"agent_name":"document-translator","search_score":0.151685216698004},{"agent_name":"warehouse-sql-analyst","search_score":0.10088637991190563}]
    
  3. curl -s 'http://localhost:$A/v1/agents?query=book+a+trip&top_k=3' -H "Authorization: Bearer $KEY_ALL" | jq -c '[.[] | {agent_name, search_score}]'
    [{"agent_name":"trip-planner","search_score":0.6537126920899458},{"agent_name":"document-translator","search_score":0.15168268706219803},{"agent_name":"warehouse-sql-analyst","search_score":0.10085822210487842}]
    
  4. curl -s 'http://localhost:$A/spend/logs?api_key=$KEY_ALL_HASH' -H 'Authorization: Bearer $MASTER' | jq -c 'sort_by(.startTime) | .[] | {model, call_type, prompt_tokens, spend}'
    {"model":"openai/text-embedding-3-small","call_type":"aembedding","prompt_tokens":3,"spend":6.000000000000001E-8}
    {"model":"openai/text-embedding-3-small","call_type":"aembedding","prompt_tokens":3,"spend":6.000000000000001E-8}
    

Dependent paths the change touches, same tip

Back on the two-deployment agent-embedder group from the section above, at 4026aa6: six searches across both vector sizes, agent_search over /mcp-rest, tools/list and tools/call over /mcp/, and a restricted key

  1. curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=plan a vacation itinerary' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'
    [{"agent_name":"trip-planner","search_score":0.42980679079763334}]
    
  2. curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=translate a pdf document' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'
    [{"agent_name":"document-translator","search_score":0.6732450375047712}]
    
  3. curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=check how many units are left in stock' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'
    [{"agent_name":"warehouse-sql-analyst","search_score":0.3791131959166272}]
    
  4. curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=book flights and hotels' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'
    [{"agent_name":"trip-planner","search_score":0.6201761715746669}]
    
  5. curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=convert a word file to spanish' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'
    [{"agent_name":"document-translator","search_score":0.5382262072991886}]
    
  6. curl -s 'http://localhost:$A/v1/agents?top_k=1' --get --data-urlencode 'query=inventory levels by warehouse' -H 'Authorization: Bearer $MASTER' | jq -c 'if type == "array" then [.[] | {agent_name, search_score}] else . end'
    [{"agent_name":"warehouse-sql-analyst","search_score":0.5334299390626673}]
    
  7. curl -s http://localhost:$A/mcp-rest/tools/call -H 'Authorization: Bearer $KEY_SEARCH' -H 'Content-Type: application/json' -d '{"name": "agent_search", "arguments": {"query": "check how many units are left in stock", "top_k": 2}}' | jq -c .
    {"_meta":null,"content":[{"type":"text","text":"[{\"agent_id\": \"b21b8787-8b9e-4c5f-a45c-3f5e4061d70e\", \"agent_name\": \"warehouse-sql-analyst\", \"description\": \"Answers questions about stock by running SQL queries against the inventory database\", \"skills\": [{\"name\": \"Query stock levels\", \"description\": \"Run a SQL query against the inventory database and summarize the stock levels it returns\", \"tags\": [\"sql\", \"analytics\"]}], \"score\": 0.3791485568539873}, {\"agent_id\": \"fc98c971-659a-42dc-902c-b8b8f59a50ef\", \"agent_name\": \"trip-planner\", \"description\": \"Books flights and hotels for a business trip given dates, origin and destination\", \"skills\": [{\"name\": \"Book travel\", \"description\": \"Find and book the flights and hotel for a business trip\", \"tags\": [\"travel\", \"booking\"]}], \"score\": 0.12168402727229097}]","annotations":null,"_meta":null}],"structuredContent":null,"isError":false}
    
  8. curl -s http://localhost:$A/mcp/ -H 'Authorization: Bearer $KEY_SEARCH' -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' -d '{"jsonrpc": "2.0", "id": 1, "method": "tools/list"}' | grep '^data:' | sed 's/^data: //' | jq -c '[.result.tools[].name]'
    ["mcp_tool_search","mcp_tool_call","agent_search"]
    
  9. curl -s http://localhost:$A/mcp/ -H 'Authorization: Bearer $KEY_SEARCH' -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' -d '{"jsonrpc": "2.0", "id": 2, "method": "tools/call", "params": {"name": "agent_search", "arguments": {"query": "book a flight and hotel", "top_k": 1}}}' | grep '^data:' | sed 's/^data: //' | jq -c '.result.content[0].text | fromjson? // .'
    [{"agent_id":"fc98c971-659a-42dc-902c-b8b8f59a50ef","agent_name":"trip-planner","description":"Books flights and hotels for a business trip given dates, origin and destination","skills":[{"name":"Book travel","description":"Find and book the flights and hotel for a business trip","tags":["travel","booking"]}],"score":0.6081681268257517}]
    
  10. curl -s 'http://localhost:$A/v1/agents?query=plan+a+vacation+itinerary&top_k=3' -H 'Authorization: Bearer $KEY_RESTRICTED' | jq -c '.[] | {agent_name, search_score}'
{"agent_name":"warehouse-sql-analyst","search_score":0.09470067778779014}

Type

🆕 New Feature

Caveats (if any)

Low

  • Keys with mcp_tool_search_enabled now see three virtual tools, not two; gating agent_search behind its own flag would add a setting nobody asked for
  • Vector cache is per process and per embedding model; each worker embeds agents once per model (again after a vector-size change), and that fill is billed to whichever key searched first on that worker
  • Over /mcp-rest, which never validates tool arguments, agent_search with no query returns the embedding provider's empty-input error as an isError result; /mcp/ rejects it up front via jsonschema
  • Unknown litellm_settings keys are accepted silently, so typos mean 400s at query time

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

  • 4026aa6 passes /live-pr-risk


Note

Medium Risk
Search triggers billable embedding calls and merges a global in-process vector cache, but behavior is scoped to agent discovery, reuses existing key RBAC, and fails explicitly when misconfigured.

Overview
Adds semantic ranking over the A2A agent registry so callers can describe a task in natural language instead of scanning every agent card.

REST: GET /v1/agents now accepts optional query and top_k. When query is set, accessible agents are embedded and ranked by cosine similarity; each row gets a search_score. Missing litellm_settings.agent_search_embedding_model returns 400; embedding failures return 503. Listing without query is unchanged (scores stay null).

Config: New agent_search_embedding_model selects an embedding model from model_list. Embedding spend is attributed to the calling API key via router metadata.

Shared core: New agent_search module implements text extraction from agent cards, router-backed embeddings, a per-model vector cache (handles mixed vector dimensions and concurrent searches), and search_agents() used by both REST and MCP.

MCP: Virtual tool agent_search joins mcp_tool_search / mcp_tool_call behind VIRTUAL_TOOL_NAMES, still gated on mcp_tool_search_enabled. REST and protocol MCP paths dispatch to the same handler and return ranked JSON (or tool errors when not configured).

RBAC: Agent listing and search use shared accessible_agents() so results only include agents the key (and team grants) can reach.

OpenAPI / dashboard schema.d.ts are updated for the new query params, search_score, and virtual tool schema typing.

Reviewed by Cursor Bugbot for commit 4026aa6. Bugbot is set up for automated code reviews on this repo. Configure here.

@greptile-apps

greptile-apps Bot commented Aug 28, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

Adds semantic ranking to the accessible A2A agent registry and exposes it through REST and MCP while attributing embedding spend to the calling key.

  • Adds query and top_k support to GET /v1/agents.
  • Introduces a concurrent, per-model agent-vector cache with mixed-dimension recovery.
  • Adds the agent_search MCP virtual tool and updates generated API types.
  • Extends tests for ranking, permissions, caching, configuration errors, and MCP dispatch.

Confidence Score: 5/5

The PR appears safe to merge because no blocking failure remains.

No blocking failure remains.

Important Files Changed

Filename Overview
litellm/proxy/agent_endpoints/agent_search.py Implements embedding-backed ranking, spend metadata propagation, dimension recovery, and concurrency-safe cache merging.
litellm/proxy/agent_endpoints/endpoints.py Adds validated search parameters and maps search configuration or embedding failures to explicit HTTP responses.
litellm/proxy/agent_endpoints/auth/agent_permission_handler.py Extracts the existing accessible-agent filtering logic for reuse by listing and search.
litellm/proxy/_experimental/mcp_server/tool_search.py Defines and dispatches the agent_search virtual tool through the shared agent-search implementation.
litellm/proxy/_experimental/mcp_server/rest_endpoints.py Routes agent_search calls from the MCP REST endpoint through the virtual-tool path.
litellm/proxy/_experimental/mcp_server/server.py Exposes and dispatches agent_search through the protocol MCP server with the existing virtual-tool permission gate.
tests/test_litellm/proxy/agent_endpoints/test_agent_search.py Covers ranking, cache concurrency, mixed vector dimensions, permission filtering, spend metadata, and error outcomes.
tests/test_litellm/proxy/_experimental/mcp_server/test_mcp_tool_search.py Covers virtual-tool schema serialization, discovery, authorization, and agent-search dispatch.

Reviews (5): Last reviewed commit: "fix(a2a): merge fresh agent vectors into..." | Re-trigger Greptile

@codecov

codecov Bot commented Aug 28, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.74775% with 5 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
...ellm/proxy/_experimental/mcp_server/tool_search.py 95.45% 2 Missing ⚠️
litellm/proxy/agent_endpoints/endpoints.py 92.59% 2 Missing ⚠️
litellm/proxy/agent_endpoints/agent_search.py 99.23% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

Comment thread litellm/proxy/agent_endpoints/agent_search.py Outdated
@veria-ai

veria-ai Bot commented Aug 28, 2026 •

Copy link
Copy Markdown
Contributor

PR overview

All previously flagged issues have been addressed. No open security concerns remain on this pull request.

Security review

No open security issues remain on this pull request.

Fixed/addressed: 1 · PR risk: 0/10

@mateo-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.

Autofix Details

Bugbot Autofix prepared a fix for the issue found in the latest run.

  • ✅ Fixed: Stale embeddings crash agent ranking
    • After embedding, the index now drops cached agent vectors whose dimension no longer matches the query and re-embeds those texts, so a dimension change from a router fallback or config reload can no longer poison the cache with mismatched vectors that make cosine_similarity raise ValueError.

Create PR

Or push these changes by commenting:

@cursor push 2ed2486bf4
Preview (2ed2486bf4)
diff --git a/litellm/proxy/agent_endpoints/agent_search.py b/litellm/proxy/agent_endpoints/agent_search.py
--- a/litellm/proxy/agent_endpoints/agent_search.py
+++ b/litellm/proxy/agent_endpoints/agent_search.py
@@ -164,10 +164,24 @@
             return AgentSearchEmbeddingFailed(
                 reason=f"embedding model returned {len(vectors)} vectors for {len(missing) + 1} inputs"
             )
-        self._vectors = MappingProxyType(dict(chain(self._vectors.items(), zip(missing, vectors[1:], strict=True))))
+        query_vector: Final = vectors[0]
+        fresh: Final = dict(zip(missing, vectors[1:], strict=True))
+        kept: Final = {text: vec for text, vec in self._vectors.items() if len(vec) == len(query_vector)}
+        stale: Final = tuple(dict.fromkeys(text for text in texts if text not in kept and text not in fresh))
+        try:
+            refreshed: Final = await embed(stale) if stale else ()
+        except (OpenAIError, ValueError, BudgetExceededError) as exc:
+            return AgentSearchEmbeddingFailed(reason=f"re-embedding stale agent texts failed: {exc}")
+        if len(refreshed) != len(stale):
+            return AgentSearchEmbeddingFailed(
+                reason=f"embedding model returned {len(refreshed)} vectors for {len(stale)} inputs"
+            )
+        self._vectors = MappingProxyType(
+            dict(chain(kept.items(), fresh.items(), zip(stale, refreshed, strict=True)))
+        )
         ranked: Final = sorted(
             (
-                AgentSearchHit(agent=agent, score=cosine_similarity(vectors[0], self._vectors[text]))
+                AgentSearchHit(agent=agent, score=cosine_similarity(query_vector, self._vectors[text]))
                 for agent, text in zip(agents, texts, strict=True)
             ),
             key=lambda hit: hit.score,

diff --git a/tests/test_litellm/proxy/agent_endpoints/test_agent_search.py b/tests/test_litellm/proxy/agent_endpoints/test_agent_search.py
--- a/tests/test_litellm/proxy/agent_endpoints/test_agent_search.py
+++ b/tests/test_litellm/proxy/agent_endpoints/test_agent_search.py
@@ -147,7 +147,20 @@
         outcome = await AgentSearchIndex().search("q", AGENTS, top_k=5, embed=short)
         assert isinstance(outcome, AgentSearchEmbeddingFailed)
 
+    @pytest.mark.asyncio
+    async def test_dimension_change_invalidates_cached_agent_vectors(self) -> None:
+        index = AgentSearchIndex()
+        warm = FakeEmbedder()
+        await index.search("language translation", AGENTS, top_k=5, embed=warm)
 
+        async def wider(texts: Sequence[str]) -> Sequence[Vector]:
+            return tuple((1.0, 0.0, 0.0, 0.0) for _ in texts)
+
+        outcome = await index.search("language translation", AGENTS, top_k=5, embed=wider)
+        assert isinstance(outcome, AgentSearchHits)
+        assert len(outcome.hits) == len(AGENTS)
+
+
 class TestSearchAgents:
     @pytest.mark.asyncio
     async def test_no_embedding_model_is_not_configured(self) -> None:

You can send follow-ups to the cloud agent here.

Comment thread litellm/proxy/agent_endpoints/agent_search.py
@mateo-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.

Autofix Details

Bugbot Autofix prepared a fix for the issue found in the latest run.

  • ✅ Fixed: Cache keeps mixed-dimension agent vectors
    • Rebuilt the cache write to re-read the current per-model map and filter both it and the embed result down to entries matching the new query dimension, so a subset re-embed evicts old-size vectors instead of merging them back from a stale pre-await snapshot.

Create PR

Or push these changes by commenting:

@cursor push efbeae4393
Preview (efbeae4393)
diff --git a/litellm/proxy/agent_endpoints/agent_search.py b/litellm/proxy/agent_endpoints/agent_search.py
--- a/litellm/proxy/agent_endpoints/agent_search.py
+++ b/litellm/proxy/agent_endpoints/agent_search.py
@@ -199,17 +199,21 @@
         if not agents:
             return AgentSearchHits(hits=())
         texts: Final = tuple(agent_search_text(agent) for agent in agents)
-        cached: Final = self._vectors.get(embedding_model, _NO_VECTORS)
-        embedded: Final = await _embed_query_and_agents(embed, query, texts, cached)
+        embedded: Final = await _embed_query_and_agents(
+            embed, query, texts, self._vectors.get(embedding_model, _NO_VECTORS)
+        )
         if isinstance(embedded, AgentSearchEmbeddingFailed):
             return embedded
         if not _same_dimension(embedded.query_vector, embedded.vectors, texts):
             return AgentSearchEmbeddingFailed(
                 reason=f"embedding model {embedding_model} returned vectors of mixed dimensions"
             )
-        self._vectors = MappingProxyType(
-            {**self._vectors, embedding_model: MappingProxyType({**cached, **embedded.vectors})}
+        current: Final = self._vectors.get(embedding_model, _NO_VECTORS)
+        dim: Final = len(embedded.query_vector)
+        merged: Final = MappingProxyType(
+            {text: vector for text, vector in chain(current.items(), embedded.vectors.items()) if len(vector) == dim}
         )
+        self._vectors = MappingProxyType({**self._vectors, embedding_model: merged})
         ranked: Final = sorted(
             (
                 AgentSearchHit(agent=agent, score=cosine_similarity(embedded.query_vector, embedded.vectors[text]))

diff --git a/tests/test_litellm/proxy/agent_endpoints/test_agent_search.py b/tests/test_litellm/proxy/agent_endpoints/test_agent_search.py
--- a/tests/test_litellm/proxy/agent_endpoints/test_agent_search.py
+++ b/tests/test_litellm/proxy/agent_endpoints/test_agent_search.py
@@ -157,6 +157,19 @@
         ]
 
     @pytest.mark.asyncio
+    async def test_subset_reembed_after_dimension_change_evicts_stale_vectors(self) -> None:
+        index = AgentSearchIndex()
+        await index.search("language translation", AGENTS, top_k=5, embed=FakeEmbedder(), embedding_model="m")
+        narrow = FixedDimensionEmbedder(2)
+        await index.search("language translation", (TRANSLATOR,), top_k=5, embed=narrow, embedding_model="m")
+        broader = FixedDimensionEmbedder(2)
+        outcome = await index.search("language translation", AGENTS, top_k=5, embed=broader, embedding_model="m")
+        assert isinstance(outcome, AgentSearchHits)
+        assert broader.calls == [
+            ("language translation", agent_search_text(SQL_ANALYST), agent_search_text(TRIP_PLANNER)),
+        ]
+
+    @pytest.mark.asyncio
     async def test_mixed_dimensions_in_one_batch_become_embedding_failed(self) -> None:
         async def mixed(texts: Sequence[str]) -> Sequence[Vector]:
             return ((1.0, 0.0), *((1.0, 0.0, 0.0) for _ in texts[1:]))

You can send follow-ups to the cloud agent here.

Comment thread litellm/proxy/agent_endpoints/agent_search.py Outdated
@codspeed

codspeed Bot commented Aug 28, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_a2a_agent_semantic_search (4026aa6) with litellm_internal_staging (98c5233)1

Open in CodSpeed

Footnotes

  1. No successful run was found on litellm_internal_staging (3300fc3) during the generation of this report, so 98c5233 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

@mateo-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 4026aa6. Configure here.

@mateo-berri
mateo-berri merged commit e1cc96e into litellm_internal_staging Aug 28, 2026
118 of 127 checks passed
@mateo-berri
mateo-berri deleted the litellm_a2a_agent_semantic_search branch August 28, 2026 17:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Feature] Semantic Agent Discovery — Search API + MCP Tool

2 participants