Skip to content

fix(moonshot, together_ai): send the reasoning effort Kimi K3 accepts - #38611

Merged
tin-berri merged 2 commits into
litellm_internal_stagingfrom
litellm_provider_effort_plumbing
Aug 28, 2026
Merged

tin-berri merged 2 commits into
litellm_internal_stagingfrom
litellm_provider_effort_plumbing

Conversation

@tin-berri

@tin-berri tin-berri commented Aug 28, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Moonshot 400s on any explicit reasoning_effort
  • Together silently folds Kimi K3 max down to high
  • Both providers document and accept low, high, max

How it solves it:

  • Moonshot offers reasoning_effort when the registry says the model reasons
  • Together sends a level the map entry declares, unchanged
  • Undeclared levels keep the existing per-model clamp
  • A bridge-derived {effort, summary} object is unwrapped to its level string

User Flow

Before: a developer routing Kimi K3 through the gateway cannot ask for max thinking, on either provider

  1. They send POST https://litellm-domain/v1/chat/completions with "model": "kimi-k3-moonshot" and "reasoning_effort": "max"
  2. It comes back 400 with moonshot does not support parameters: ['reasoning_effort'], and the request never left the gateway
  3. They drop the parameter to get a 200, and every call now runs at whatever effort the provider defaults to
  4. They try the same thing against Together with "model": "kimi-k3-together" and get a 200, so it looks fine
  5. Together bills and answers at high, because the gateway rewrote max on the way out, and nothing in the response says so
  6. Their Anthropic-format client sends POST https://litellm-domain/v1/messages with a thinking block and gets the same 400

After: both providers take the level the caller asked for

  1. They send the same POST with "model": "kimi-k3-moonshot" and "reasoning_effort": "max"
  2. It comes back 200 with a completion
  3. They send the same POST with "model": "kimi-k3-together" and "reasoning_effort": "max"
  4. It comes back 200, and Together receives "reasoning_effort": "max" rather than high
  5. Asking for low or high on either provider does what it says, and both are reachable for the first time on Moonshot
  6. The same POST https://litellm-domain/v1/messages with a thinking block comes back 200 with a thinking content block

Relevant issues

  • Moonshot never listed reasoning_effort among its supported params, so an explicit level raised before the request left the proxy
  • Together clamps max to high for every model except one prefix, which caught Kimi K3 after the map declared it takes max
  • Only Kimi K3 on Together changes value; every other Together model keeps its current clamp exactly
  • The /v1/messages and /v1/responses bridges already derived reasoning_effort for non-Claude targets, so the same 400 blocked both surfaces; when a thinking summary is in play the derived value is an object like {"effort": "high", "summary": "detailed"}, which Moonshot rejects with "expected type string" even once the param is allowed, so the config unwraps it to the level string

Linear ticket

Resolves LIT-6330

Pre-Submission checklist

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review

Screenshots / Proof of Fix

Shared setup, a proxy with both providers registered on their real APIs, launched with LITELLM_LOCAL_MODEL_COST_MAP=true so the bundled map's reasoning_effort_levels entry is loaded (see Caveats):

model_list:
  - model_name: kimi-k3-moonshot
    litellm_params:
      model: moonshot/kimi-k3
      api_key: os.environ/MOONSHOT_API_KEY
  - model_name: kimi-k3-together
    litellm_params:
      model: together_ai/moonshotai/Kimi-K3
      api_key: os.environ/TOGETHERAI_API_KEY

Outgoing bodies are captured by pointing one extra deployment per provider at a local listener, since a downgrade is invisible from the client side.

Before (74050e0)

/v1/chat/completions, Moonshot rejects the parameter

  1. curl -s -X POST http://127.0.0.1:16422/v1/chat/completions -H "Authorization: Bearer sk-1234" -d '{"model":"kimi-k3-moonshot","messages":[{"role":"user","content":"Reply ok."}],"reasoning_effort":"max","max_tokens":2048}'
  2. status=400, body litellm.UnsupportedParamsError: moonshot does not support parameters: ['reasoning_effort'], for model=kimi-k3

/v1/chat/completions, Together downgrades max on the wire

  1. Same request against the wire-capture deployment for each of low, high, max, minimal, xhigh

  2. What Together received:

    asked sent
    low low
    high high
    max high
    minimal low
    xhigh high

/v1/messages, thinking never reaches Moonshot

  1. curl -s -X POST http://127.0.0.1:16422/v1/messages -H "Authorization: Bearer sk-1234" -d '{"model":"kimi-k3-moonshot","max_tokens":4096,"messages":[{"role":"user","content":"Reply ok."}],"thinking":{"type":"enabled","budget_tokens":4096}}'
  2. status=400, same UnsupportedParamsError for ['reasoning_effort']

/v1/responses, reasoning effort never reaches Moonshot

  1. curl -s -X POST http://127.0.0.1:16422/v1/responses -H "Authorization: Bearer sk-1234" -d '{"model":"kimi-k3-moonshot","input":"Reply ok.","reasoning":{"effort":"max"}}'
  2. status=400, same UnsupportedParamsError for ['reasoning_effort']

After (7fd136d)

/v1/chat/completions, Moonshot rejects the parameter

  1. curl -s -X POST http://127.0.0.1:16431/v1/chat/completions -H "Authorization: Bearer sk-1234" -d '{"model":"kimi-k3-moonshot","messages":[{"role":"user","content":"Reply ok."}],"reasoning_effort":"max","max_tokens":2048}'
  2. status=200, model=kimi-k3-moonshot completion_tokens=46

/v1/chat/completions, Together downgrades max on the wire

  1. Same five requests against the wire-capture deployment

  2. What Together received:

    asked before after
    low low low
    high high high
    max high max
    minimal low low
    xhigh high high
  3. A real call to kimi-k3-together at max returns status=200

/v1/messages, thinking never reaches Moonshot

  1. Same request as before, status=200, and the wire shows reasoning_effort: "high" derived from the 4096 budget
  2. With "thinking":{"type":"enabled","budget_tokens":8192,"summary":"detailed"} the wire still shows the plain string "high" rather than the {"effort","summary"} object, and the real call returns status=200 with content types ['thinking', 'text']; before the unwrap commit Moonshot answered that object with 400 the reasoning_effort field (expected type string) is illegal

/v1/responses, reasoning effort never reaches Moonshot

  1. Same request as before, status=200, wire shows reasoning_effort: "max"

Why max is worth having

Measured directly against both provider APIs on one prompt, so the downgrade is not cosmetic:

provider high max
Moonshot kimi-k3 400 reasoning tokens 872
Together Kimi-K3 453 reasoning tokens 889

Type

🐛 Bug Fix

Caveats (if any)

Medium

  • azure_ai also rejects the param, left out because Foundry is unverified
  • LIT-6330 covers azure_ai too, so it stays open
  • Together passthrough reads reasoning_effort_levels off the loaded map
    • the public remote map carries the key on no entry yet
    • a default proxy therefore keeps the max to high clamp until it ships
    • LITELLM_LOCAL_MODEL_COST_MAP=true or a model_info declaration activates it today

Low

  • Bedrock Moonshot skips this config's param list, so it is unaffected
  • 5 of 23 Moonshot entries gain the param, all kimi reasoning models
  • Together tests copy the clamp tables' literals rather than importing them
  • A bridge thinking summary has no Moonshot equivalent and is dropped

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Note

Medium Risk
Changes request-parameter mapping on the hot path for Moonshot and Together chat completions; scope is narrow and well-tested, but wrong mapping could still change provider billing or reasoning behavior for Kimi reasoning models.

Overview
Fixes Moonshot rejecting reasoning_effort and Together silently rewriting max to high for Kimi K3-style models that declare low, high, and max in the model map.

Moonshot now advertises reasoning_effort for models where supports_reasoning is true, maps it in map_openai_params, and unwraps bridge-style {"effort", "summary"} values to the bare string Moonshot expects (invalid or missing effort strings are omitted).

Together checks declared_reasoning_efforts_for_model before the hardcoded effort translation tables, so declared levels like max are sent unchanged; undeclared levels and none still follow the existing clamps and disable-reasoning behavior.

Adds declared_reasoning_efforts_for_model in reasoning_effort_capability.py to read reasoning_effort_levels from model_cost without raising on unknown models. Unit tests cover Moonshot optional-params and unwrap behavior and Together declared vs undeclared effort mapping.

Reviewed by Cursor Bugbot for commit 7fd136d. Bugbot is set up for automated code reviews on this repo. Configure here.

Moonshot documents reasoning_effort as a top-level chat completions field for its reasoning
models, and defaults it to max, but MoonshotChatConfig builds its supported params by
subtracting from the OpenAI base list, which never carried that param. An explicit level
raised UnsupportedParamsError before the request left the proxy, so low and high were
unreachable and every call ran at the provider default

Together accepts low, high and max on Kimi K3. The per-model clamp added for the gpt-oss
family folds max down to high for every model except deepseek-ai/DeepSeek-V4-Pro, so a caller
asking for max silently got roughly half the reasoning budget they paid for

Moonshot now offers reasoning_effort whenever the registry says the model reasons. Together
sends a level the map entry declares unchanged, and keeps its existing table for every level
an entry does not name, so the only value that moves is Kimi K3 at max
@tin-berri

Copy link
Copy Markdown
Contributor Author

Same fix shape you all landed on, now with live proof on both providers.

@greptileai can you reply here with a score and why that score?

bugbot run

@greptile-apps

greptile-apps Bot commented Aug 28, 2026

Copy link
Copy Markdown
Contributor

Confidence: 5/5.

The fix is correctly scoped and addresses both regressions:

  • litellm/llms/moonshot/chat/transformation.py:106 now advertises reasoning_effort only when the registry marks the model as reasoning-capable, while preserving rejection for non-reasoning models.
  • litellm/llms/together_ai/chat/transformation.py:145 gives registry-declared levels precedence, so Kimi K3's low, high, and max values are sent unchanged while undeclared values retain the existing clamp behavior and none still disables reasoning.
  • litellm/router_utils/reasoning_effort_capability.py:91 safely handles provider-prefixed and unmapped model names without turning lookup misses into request failures.
  • The added tests cover supported and unsupported Moonshot models, all declared Kimi K3 levels, undeclared-level regression behavior, none, the public optional-params path, and the live-provider results documented in the PR.

I found no correctness or regression concerns, and the reported CI status is passing.

@greptile-apps

greptile-apps Bot commented Aug 28, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR forwards supported reasoning effort values without changing existing fallback translations

  • Moonshot advertises reasoning_effort for models marked as reasoning-capable
  • Together preserves effort levels explicitly declared by model metadata
  • Shared capability lookup and focused regression tests cover low, high, max, undeclared levels, and disabled reasoning

Confidence Score: 5/5

The PR appears safe to merge, with no concrete blocking or non-blocking defects identified

The changed adapters preserve declared reasoning levels while retaining existing behavior for undeclared values, unsupported models, default effort, and disabled reasoning

Important Files Changed

Filename Overview
litellm/llms/moonshot/chat/transformation.py Adds reasoning_effort to the supported parameter list for Moonshot reasoning models while retaining existing exclusions
litellm/llms/together_ai/chat/transformation.py Preserves model-declared reasoning effort values before applying the existing Together fallback translations
litellm/router_utils/reasoning_effort_capability.py Adds a non-raising model catalog lookup for provider adapters that need declared reasoning effort levels
tests/test_litellm/llms/moonshot/test_moonshot_chat_transformation.py Covers Moonshot capability advertisement, parameter forwarding, and rejection for non-reasoning models
tests/test_litellm/llms/together_ai/chat/test_together_ai_chat_transformation.py Covers declared-level passthrough and unchanged fallback behavior for undeclared Together effort values

Reviews (1): Last reviewed commit: "fix(moonshot, together_ai): send the rea..." | Re-trigger Greptile

@codecov

codecov Bot commented Aug 28, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 95.45455% with 1 line in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
...itellm/router_utils/reasoning_effort_capability.py 80.00% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 3412352. Configure here.

@codspeed

codspeed Bot commented Aug 28, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_provider_effort_plumbing (7fd136d) with litellm_internal_staging (98c5233)

Open in CodSpeed

@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai can you reply here with score & why

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Aug 28, 2026

Copy link
Copy Markdown
Contributor

Confidence: 4/5

The implementation is focused and the behavior is well covered:

  • Moonshot now advertises reasoning_effort only for registry-backed reasoning models, while non-reasoning models retain the existing rejection behavior.
  • Moonshot safely unwraps bridge-derived {effort, summary} values to the provider’s required level string and omits malformed values.
  • Together passes declared Kimi K3 levels (low, high, max) unchanged, while preserving the existing mappings for undeclared levels and none.
  • Tests cover the supported levels, legacy clamps, bridge-object unwrapping, invalid values, and optional-parameter integration.
  • The available CI status is successful.

I’m holding back one point because the Together fix is data-dependent: as noted in the PR, the current public remote model map does not yet declare Kimi K3’s reasoning_effort_levels, so default deployments may continue clamping max until that registry data ships or the model is configured locally.

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 7fd136d. Configure here.

@tin-berri
tin-berri enabled auto-merge (squash) August 28, 2026 02:56
@tin-berri
tin-berri disabled auto-merge August 28, 2026 03:46
@tin-berri
tin-berri merged commit 3300fc3 into litellm_internal_staging Aug 28, 2026
82 checks passed
@tin-berri
tin-berri deleted the litellm_provider_effort_plumbing branch August 28, 2026 03:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants