fix(moonshot, together_ai): send the reasoning effort Kimi K3 accepts - #38611
Conversation
Moonshot documents reasoning_effort as a top-level chat completions field for its reasoning models, and defaults it to max, but MoonshotChatConfig builds its supported params by subtracting from the OpenAI base list, which never carried that param. An explicit level raised UnsupportedParamsError before the request left the proxy, so low and high were unreachable and every call ran at the provider default Together accepts low, high and max on Kimi K3. The per-model clamp added for the gpt-oss family folds max down to high for every model except deepseek-ai/DeepSeek-V4-Pro, so a caller asking for max silently got roughly half the reasoning budget they paid for Moonshot now offers reasoning_effort whenever the registry says the model reasons. Together sends a level the map entry declares unchanged, and keeps its existing table for every level an entry does not name, so the only value that moves is Kimi K3 at max
|
Same fix shape you all landed on, now with live proof on both providers. @greptileai can you reply here with a score and why that score? bugbot run |
|
Confidence: 5/5. The fix is correctly scoped and addresses both regressions:
I found no correctness or regression concerns, and the reported CI status is passing. |
Greptile SummaryThis PR forwards supported reasoning effort values without changing existing fallback translations
Confidence Score: 5/5The PR appears safe to merge, with no concrete blocking or non-blocking defects identified The changed adapters preserve declared reasoning levels while retaining existing behavior for undeclared values, unsupported models, default effort, and disabled reasoning
|
| Filename | Overview |
|---|---|
| litellm/llms/moonshot/chat/transformation.py | Adds reasoning_effort to the supported parameter list for Moonshot reasoning models while retaining existing exclusions |
| litellm/llms/together_ai/chat/transformation.py | Preserves model-declared reasoning effort values before applying the existing Together fallback translations |
| litellm/router_utils/reasoning_effort_capability.py | Adds a non-raising model catalog lookup for provider adapters that need declared reasoning effort levels |
| tests/test_litellm/llms/moonshot/test_moonshot_chat_transformation.py | Covers Moonshot capability advertisement, parameter forwarding, and rejection for non-reasoning models |
| tests/test_litellm/llms/together_ai/chat/test_together_ai_chat_transformation.py | Covers declared-level passthrough and unchanged fallback behavior for undeclared Together effort values |
Reviews (1): Last reviewed commit: "fix(moonshot, together_ai): send the rea..." | Re-trigger Greptile
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 3412352. Configure here.
|
@greptileai can you reply here with score & why |
|
bugbot run |
|
Confidence: 4/5 The implementation is focused and the behavior is well covered:
I’m holding back one point because the Together fix is data-dependent: as noted in the PR, the current public remote model map does not yet declare Kimi K3’s |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 7fd136d. Configure here.
TLDR
Problem this solves:
How it solves it:
User Flow
Before: a developer routing Kimi K3 through the gateway cannot ask for max thinking, on either provider
"model": "kimi-k3-moonshot"and"reasoning_effort": "max"moonshot does not support parameters: ['reasoning_effort'], and the request never left the gateway"model": "kimi-k3-together"and get a 200, so it looks finethinkingblock and gets the same 400After: both providers take the level the caller asked for
"model": "kimi-k3-moonshot"and"reasoning_effort": "max""model": "kimi-k3-together"and"reasoning_effort": "max""reasoning_effort": "max"rather than highthinkingblock comes back 200 with a thinking content blockRelevant issues
reasoning_effortamong its supported params, so an explicit level raised before the request left the proxymaxtohighfor every model except one prefix, which caught Kimi K3 after the map declared it takes maxreasoning_effortfor non-Claude targets, so the same 400 blocked both surfaces; when a thinking summary is in play the derived value is an object like{"effort": "high", "summary": "detailed"}, which Moonshot rejects with "expected type string" even once the param is allowed, so the config unwraps it to the level stringLinear ticket
Resolves LIT-6330
Pre-Submission checklist
uv run pytest tests/test_litellm/<your_test_file>.py -vScreenshots / Proof of Fix
Shared setup, a proxy with both providers registered on their real APIs, launched with
LITELLM_LOCAL_MODEL_COST_MAP=trueso the bundled map'sreasoning_effort_levelsentry is loaded (see Caveats):Outgoing bodies are captured by pointing one extra deployment per provider at a local listener, since a downgrade is invisible from the client side.
Before (74050e0)
/v1/chat/completions, Moonshot rejects the parameter
curl -s -X POST http://127.0.0.1:16422/v1/chat/completions -H "Authorization: Bearer sk-1234" -d '{"model":"kimi-k3-moonshot","messages":[{"role":"user","content":"Reply ok."}],"reasoning_effort":"max","max_tokens":2048}'status=400, bodylitellm.UnsupportedParamsError: moonshot does not support parameters: ['reasoning_effort'], for model=kimi-k3/v1/chat/completions, Together downgrades max on the wire
Same request against the wire-capture deployment for each of low, high, max, minimal, xhigh
What Together received:
/v1/messages, thinking never reaches Moonshot
curl -s -X POST http://127.0.0.1:16422/v1/messages -H "Authorization: Bearer sk-1234" -d '{"model":"kimi-k3-moonshot","max_tokens":4096,"messages":[{"role":"user","content":"Reply ok."}],"thinking":{"type":"enabled","budget_tokens":4096}}'status=400, sameUnsupportedParamsErrorfor['reasoning_effort']/v1/responses, reasoning effort never reaches Moonshot
curl -s -X POST http://127.0.0.1:16422/v1/responses -H "Authorization: Bearer sk-1234" -d '{"model":"kimi-k3-moonshot","input":"Reply ok.","reasoning":{"effort":"max"}}'status=400, sameUnsupportedParamsErrorfor['reasoning_effort']After (7fd136d)
/v1/chat/completions, Moonshot rejects the parameter
curl -s -X POST http://127.0.0.1:16431/v1/chat/completions -H "Authorization: Bearer sk-1234" -d '{"model":"kimi-k3-moonshot","messages":[{"role":"user","content":"Reply ok."}],"reasoning_effort":"max","max_tokens":2048}'status=200,model=kimi-k3-moonshot completion_tokens=46/v1/chat/completions, Together downgrades max on the wire
Same five requests against the wire-capture deployment
What Together received:
A real call to
kimi-k3-togetherat max returnsstatus=200/v1/messages, thinking never reaches Moonshot
status=200, and the wire showsreasoning_effort: "high"derived from the 4096 budget"thinking":{"type":"enabled","budget_tokens":8192,"summary":"detailed"}the wire still shows the plain string"high"rather than the{"effort","summary"}object, and the real call returnsstatus=200with content types['thinking', 'text']; before the unwrap commit Moonshot answered that object with400 the reasoning_effort field (expected type string) is illegal/v1/responses, reasoning effort never reaches Moonshot
status=200, wire showsreasoning_effort: "max"Why max is worth having
Measured directly against both provider APIs on one prompt, so the downgrade is not cosmetic:
Type
🐛 Bug Fix
Caveats (if any)
Medium
reasoning_effort_levelsoff the loaded mapLITELLM_LOCAL_MODEL_COST_MAP=trueor amodel_infodeclaration activates it todayLow
Final Attestation
Note
Medium Risk
Changes request-parameter mapping on the hot path for Moonshot and Together chat completions; scope is narrow and well-tested, but wrong mapping could still change provider billing or reasoning behavior for Kimi reasoning models.
Overview
Fixes Moonshot rejecting
reasoning_effortand Together silently rewritingmaxtohighfor Kimi K3-style models that declarelow,high, andmaxin the model map.Moonshot now advertises
reasoning_effortfor models wheresupports_reasoningis true, maps it inmap_openai_params, and unwraps bridge-style{"effort", "summary"}values to the bare string Moonshot expects (invalid or missing effort strings are omitted).Together checks
declared_reasoning_efforts_for_modelbefore the hardcoded effort translation tables, so declared levels likemaxare sent unchanged; undeclared levels andnonestill follow the existing clamps and disable-reasoning behavior.Adds
declared_reasoning_efforts_for_modelinreasoning_effort_capability.pyto readreasoning_effort_levelsfrommodel_costwithout raising on unknown models. Unit tests cover Moonshot optional-params and unwrap behavior and Together declared vs undeclared effort mapping.Reviewed by Cursor Bugbot for commit 7fd136d. Bugbot is set up for automated code reviews on this repo. Configure here.