Skip to content

fix(models): stop LiteLlm from running tool calls quoted in answer text - #7367

Open
donggyun112 wants to merge 1 commit into
google:mainfrom
donggyun112:fix/litellm-inline-tool-call-json
Open

donggyun112 wants to merge 1 commit into
google:mainfrom
donggyun112:fix/litellm-inline-tool-call-json

Conversation

@donggyun112

Copy link
Copy Markdown
Contributor

Description of Change

Problem: When a LiteLLM provider returns no structured tool_calls, LiteLlm scans the answer text for any {...} object with a "name" string and an "arguments" key, and turns each one into a function call (lite_llm.py#L2067-L2118). The object can sit anywhere in the text, after any amount of prose. The check runs for every provider, on answers that finished with "stop", and on every streamed delta (#L2500, #L2690).

So an answer that only shows a tool call becomes one. Examples are an answer explaining a README example, quoting an API document, or repeating text from a retrieved web page. The agent then runs the named tool with those arguments, and the quoted text is cut out of the answer the user sees. Content the agent reads, such as a document or page, can steer which of the agent's tools runs with which arguments by getting the model to quote it.

The fallback was added for #1968, where a vLLM server without a tool-call parser returned the whole tool call as the message text ("\n{\"name\": \"get_current_time\", \"arguments\": {...}}\n"). The quoted-in-prose case is outside that.

Solution:

  • Treat text as tool calls only when the text opens with them. Leading whitespace and the wrappers models emit calls in are allowed: Hermes/Qwen-style <tool_call>…</tool_call> and a ```/```json fence. Consecutive calls are kept. Parsing stops at the first text that is not a call, and that text stays in the answer.
  • Decide only on whole text. A single streamed delta cannot tell a call from one quoted mid-answer, so deltas and partial responses are no longer parsed for calls. Streamed text is parsed once when the stream finalizes, which is already where calls split across deltas were found.
  • The wrapper markers around a parsed call are consumed with it. Before this change they stayed behind as visible text, for example <tool_call>\n\n</tool_call>.

DeepSeek's special-token format is unchanged.

Reproduction

Environment details: Current source based on fdca5e7, ADK 2.10.0 package metadata, macOS 26.6.2 arm64, Python 3.12, litellm 1.103.1.

Model information: LiteLLM: Yes, LiteLlm(model="openai/gpt-4o") with its llm_client replaced by a scripted client, so no network or API key is needed. The provider returns plain text with finish_reason="stop" and no tool_calls.

import asyncio

from google.adk.agents import LlmAgent
from google.adk.models.lite_llm import LiteLlm
from google.adk.runners import InMemoryRunner
from google.genai import types
from litellm import ModelResponse

ANSWER = (
    "The README shows this example request:\n"
    '{"name": "delete_file", "arguments": {"path": "/data/prod.db"}}\n'
    "It removes files."
)
deleted = []


class ScriptedClient:
  calls = 0

  async def acompletion(self, **kwargs):
    ScriptedClient.calls += 1
    text = ANSWER if ScriptedClient.calls == 1 else "ok"
    return ModelResponse(
        choices=[{
            "index": 0,
            "finish_reason": "stop",
            "message": {"role": "assistant", "content": text},
        }]
    )


def delete_file(path: str) -> str:
  """Deletes a file."""
  deleted.append(path)
  return "deleted"


async def main():
  llm = LiteLlm(model="openai/gpt-4o")
  llm.llm_client = ScriptedClient()
  agent = LlmAgent(name="a", model=llm, tools=[delete_file])
  runner = InMemoryRunner(agent=agent, app_name="app")
  session = await runner.session_service.create_session(
      app_name="app", user_id="u"
  )
  message = types.Content(
      role="user", parts=[types.Part(text="What does the README example do?")]
  )
  async for _ in runner.run_async(
      user_id="u", session_id=session.id, new_message=message
  ):
    pass
  print("delete_file ran with:", deleted)


asyncio.run(main())

Observed behavior on main:

delete_file ran with: ['/data/prod.db']

The user is shown 'The README shows this example request:\n\n', with the quoted JSON removed.

Expected behavior, verified after the correction:

delete_file ran with: []

The answer is returned unchanged, quoted JSON included.

Testing Plan

  • New tests failed on unchanged source and pass with the fix:

    • test_split_message_content_keeps_tool_call_json_quoted_in_text
    • test_parse_tool_calls_from_text_stops_at_text_between_calls
    • test_streaming_text_quoting_a_tool_call_is_not_a_call, where the prose and the JSON arrive as separate deltas
    • test_split_message_content_parses_text_that_is_a_tool_call[tool_call_tags|code_fence]: the call was already found, but the wrapper was left behind as text
  • Supported shapes stay covered:

  • Existing tests whose expectations change, because they asserted extraction after prose:

    • test_split_message_content_and_tool_calls_inline_text ('Intro {...} trailing content'), replaced by the quoted-in-text test above
    • test_parse_tool_calls_from_text_multiple_calls: the second call followed filler text, so it is now directly after the first. The filler case is the new ..._stops_at_text_between_calls.
    • test_parse_tool_calls_from_text_mixed_formats: the generic call now follows the DeepSeek block instead of following prose
    • test_model_response_to_chunk[...-tool_calls] ('Intro {...} wrap'): the message now opens with the call
  • Runner E2E with built baseline and fixed wheels installed into the same clean environment (litellm 1.103.1). Each scenario returns the scripted text from the provider, non-streaming and with StreamingMode.SSE:

    Provider text main: tool ran / first text shown fixed: tool ran / first text shown
    prose quoting a call yes / 'The README shows this example request:\n\n' no / the full answer
    whole message is a call ([ Question ] Google ADK Tools + VLLM Model + LiteLLM integration problem #1968) yes / 'ok' (stream: '\n\n') yes / 'ok'
    call in <tool_call> tags yes / '<tool_call>\n\n</tool_call>' yes / 'ok'
  • Full tests/unittests suite on each Python version from 3.10 through 3.14, using a locally generated uv.lock and --extra test:

    Python Passed Failed Skipped Xfailed Xpassed
    3.10.20 16,853 2 89 25 2
    3.11.14 16,862 2 88 25 2
    3.12.10 16,852 3 89 25 2
    3.13.1 16,853 2 89 25 2
    3.14.7 16,852 3 89 25 2

    None of the failures involves this change, and all of them also fail on unchanged main:

    • test_interactions_utils.py::TestBuildGenerationConfig::test_dropped_parameters_are_the_ones_the_request_cannot_carry fails with the resolved google-genai 2.26.0.
    • test_import_loading.py::test_entry_point_loads_only_allowlisted_packages[agent]
    • test_mcp_session_manager.py::TestMCPSessionManager::test_is_session_disconnected_without_streams failed on 3.12 and 3.14 under -n 8, and passes when run on its own.

    tests/unittests/integrations/livekit was excluded on every version because it hangs on current main as well as on this branch. It passes on 044a1ec, and the hang starts at 018c41c. The 3.14 row is from a uv-managed CPython 3.14.7, because Homebrew's CPython 3.14 ships a sitecustomize.py that test_import_loading.py reports on any tree. gcloud was excluded from PATH and stdin was closed for all runs.

  • Changed-file pre-commit passes.

  • Mypy with the CI procedure (uv sync --all-extras, errors with line numbers stripped, baseline vs. branch) on Python 3.12: 826 errors on each side and no new diagnostics. One pre-existing streaming_utils.py error, a file this change does not touch, appears in both runs with its Literal[...] members printed in a different order, so a plain text diff lists it once.

These are macOS arm64 runs. Linux CI's full-extras matrix was not reproduced locally.

Compatibility

No API change. A call that follows prose in the same text, which was extracted before, is now returned as text. That is the case this change targets. A model that writes a sentence and then a tool call as text is indistinguishable from one quoting a call. Servers that need text tool calls parsed can do it where the format is known, for example vLLM's --tool-call-parser, which returns structured tool_calls that this fallback never touches.

This change does not check the parsed name against the tools in the request, and it still parses text when the request declared no tools. Restricting the fallback to requests that declared tools, and to names among them, would follow the rule the Gemma adapter and LiteLLM's own Ollama transform use: only parse text tool calls you asked for. It would require passing the request's tool names into the conversion helpers. Would you like that as a follow-up, or should the fallback become opt-in per model?

When a provider returns no structured tool_calls, LiteLlm turned every
{"name": ..., "arguments": ...} object found anywhere in the answer text
into a function call, including one the model only quoted from a README,
an API document or a retrieved page. The agent then ran that tool and the
quoted text disappeared from the answer.

Treat text as tool calls only when it opens with them, allowing leading
whitespace and the <tool_call> and code-fence wrappers models emit, and
consume those wrappers with the call. Parse streamed text for calls only
once it is complete, since a single delta cannot tell a call from a quote.
@donggyun112
donggyun112 force-pushed the fix/litellm-inline-tool-call-json branch from c11d522 to ee3c518 Compare October 1, 2026 09:09
@donggyun112 donggyun112 closed this Oct 1, 2026
@donggyun112 donggyun112 reopened this Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants