fix(models): stop LiteLlm from running tool calls quoted in answer text - #7367
Open
donggyun112 wants to merge 1 commit into
Open
donggyun112 wants to merge 1 commit into
donggyun112 wants to merge 1 commit into
Conversation
When a provider returns no structured tool_calls, LiteLlm turned every
{"name": ..., "arguments": ...} object found anywhere in the answer text
into a function call, including one the model only quoted from a README,
an API document or a retrieved page. The agent then ran that tool and the
quoted text disappeared from the answer.
Treat text as tool calls only when it opens with them, allowing leading
whitespace and the <tool_call> and code-fence wrappers models emit, and
consume those wrappers with the call. Parse streamed text for calls only
once it is complete, since a single delta cannot tell a call from a quote.
donggyun112
force-pushed
the
fix/litellm-inline-tool-call-json
branch
from
October 1, 2026 09:09
c11d522 to
ee3c518
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description of Change
Problem: When a LiteLLM provider returns no structured
tool_calls,LiteLlmscans the answer text for any{...}object with a"name"string and an"arguments"key, and turns each one into a function call (lite_llm.py#L2067-L2118). The object can sit anywhere in the text, after any amount of prose. The check runs for every provider, on answers that finished with"stop", and on every streamed delta (#L2500, #L2690).So an answer that only shows a tool call becomes one. Examples are an answer explaining a README example, quoting an API document, or repeating text from a retrieved web page. The agent then runs the named tool with those arguments, and the quoted text is cut out of the answer the user sees. Content the agent reads, such as a document or page, can steer which of the agent's tools runs with which arguments by getting the model to quote it.
The fallback was added for #1968, where a vLLM server without a tool-call parser returned the whole tool call as the message text (
"\n{\"name\": \"get_current_time\", \"arguments\": {...}}\n"). The quoted-in-prose case is outside that.Solution:
<tool_call>…</tool_call>and a```/```jsonfence. Consecutive calls are kept. Parsing stops at the first text that is not a call, and that text stays in the answer.<tool_call>\n\n</tool_call>.DeepSeek's special-token format is unchanged.
Reproduction
Environment details: Current source based on
fdca5e7, ADK 2.10.0 package metadata, macOS 26.6.2 arm64, Python 3.12, litellm 1.103.1.Model information: LiteLLM: Yes,
LiteLlm(model="openai/gpt-4o")with itsllm_clientreplaced by a scripted client, so no network or API key is needed. The provider returns plain text withfinish_reason="stop"and notool_calls.Observed behavior on main:
The user is shown
'The README shows this example request:\n\n', with the quoted JSON removed.Expected behavior, verified after the correction:
The answer is returned unchanged, quoted JSON included.
Testing Plan
New tests failed on unchanged source and pass with the fix:
test_split_message_content_keeps_tool_call_json_quoted_in_texttest_parse_tool_calls_from_text_stops_at_text_between_callstest_streaming_text_quoting_a_tool_call_is_not_a_call, where the prose and the JSON arrive as separate deltastest_split_message_content_parses_text_that_is_a_tool_call[tool_call_tags|code_fence]: the call was already found, but the wrapper was left behind as textSupported shapes stay covered:
[bare], the [ Question ] Google ADK Tools + VLLM Model + LiteLLM integration problem #1968 shapetest_streaming_text_that_is_a_tool_call_becomes_a_call, with the call split across deltastest_message_to_generate_content_response_inline_tool_call_text, where trailing text such as<|im_end|>systemstays as textExisting tests whose expectations change, because they asserted extraction after prose:
test_split_message_content_and_tool_calls_inline_text('Intro {...} trailing content'), replaced by the quoted-in-text test abovetest_parse_tool_calls_from_text_multiple_calls: the second call followed filler text, so it is now directly after the first. The filler case is the new..._stops_at_text_between_calls.test_parse_tool_calls_from_text_mixed_formats: the generic call now follows the DeepSeek block instead of following prosetest_model_response_to_chunk[...-tool_calls]('Intro {...} wrap'): the message now opens with the callRunner E2E with built baseline and fixed wheels installed into the same clean environment (litellm 1.103.1). Each scenario returns the scripted text from the provider, non-streaming and with
StreamingMode.SSE:'The README shows this example request:\n\n''ok'(stream:'\n\n')'ok'<tool_call>tags'<tool_call>\n\n</tool_call>''ok'Full
tests/unittestssuite on each Python version from 3.10 through 3.14, using a locally generateduv.lockand--extra test:None of the failures involves this change, and all of them also fail on unchanged main:
test_interactions_utils.py::TestBuildGenerationConfig::test_dropped_parameters_are_the_ones_the_request_cannot_carryfails with the resolvedgoogle-genai2.26.0.test_import_loading.py::test_entry_point_loads_only_allowlisted_packages[agent]test_mcp_session_manager.py::TestMCPSessionManager::test_is_session_disconnected_without_streamsfailed on 3.12 and 3.14 under-n 8, and passes when run on its own.tests/unittests/integrations/livekitwas excluded on every version because it hangs on current main as well as on this branch. It passes on044a1ec, and the hang starts at 018c41c. The 3.14 row is from a uv-managed CPython 3.14.7, because Homebrew's CPython 3.14 ships asitecustomize.pythattest_import_loading.pyreports on any tree. gcloud was excluded from PATH and stdin was closed for all runs.Changed-file
pre-commitpasses.Mypy with the CI procedure (
uv sync --all-extras, errors with line numbers stripped, baseline vs. branch) on Python 3.12: 826 errors on each side and no new diagnostics. One pre-existingstreaming_utils.pyerror, a file this change does not touch, appears in both runs with itsLiteral[...]members printed in a different order, so a plain text diff lists it once.These are macOS arm64 runs. Linux CI's full-extras matrix was not reproduced locally.
Compatibility
No API change. A call that follows prose in the same text, which was extracted before, is now returned as text. That is the case this change targets. A model that writes a sentence and then a tool call as text is indistinguishable from one quoting a call. Servers that need text tool calls parsed can do it where the format is known, for example vLLM's
--tool-call-parser, which returns structuredtool_callsthat this fallback never touches.This change does not check the parsed name against the tools in the request, and it still parses text when the request declared no tools. Restricting the fallback to requests that declared tools, and to names among them, would follow the rule the Gemma adapter and LiteLLM's own Ollama transform use: only parse text tool calls you asked for. It would require passing the request's tool names into the conversion helpers. Would you like that as a follow-up, or should the fallback become opt-in per model?