Skip to content

LiteLlm drops GenerateContentConfig.thinking_config; map it to litellm's reasoning_effort #7360

Description

@chriswithers-fuse

google.adk.models.lite_llm.LiteLlm._get_completion_inputs builds the litellm kwargs from a fixed allow-list of GenerateContentConfig fields (temperature, max_output_tokens, top_p, top_k, seed, stop_sequences, presence_penalty, frequency_penalty). thinking_config is present in config_dict but never forwarded, so an agent that sets e.g. ThinkingConfig(thinking_level=ThinkingLevel.LOW) gets the backend's default reasoning effort when run through LiteLlm, with no warning.

Verified on 2.8.0 and 2.10.0 (src/google/adk/models/lite_llm.py, the param_mapping / for key in (...) block, around line 3024 in v2.10.0). The only thinking handling in the file is the Anthropic thinking-block parsing on the response side.

litellm already normalises this across providers: reasoning_effort (low / medium / high, translated per provider, e.g. to Gemini thinkingConfig), and thinking={"type": "enabled", "budget_tokens": N} for budget-style providers.

Proposal: in _get_completion_inputs, map

  • thinking_config.thinking_level → reasoning_effort
  • thinking_config.thinking_budget → thinking={"type": "enabled", "budget_tokens": ...}
  • thinking_config.include_thoughts → the provider's equivalent where litellm exposes one

Related but different: #5712 (thinking blocks dropped when parsing responses), #5805 (thinking_config ignored in run_live).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

models[Component] This issue is related to model support

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions