fix(fireworks_ai): flatten dict-form reasoning_effort to its effort string - #41335
Conversation
…tring Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
🤖 Devin AI EngineerI'll be helping with this pull request! Here's what you should know: ✅ I will automatically:
Note: I can only respond to comments from users who have write access to this repository. ⚙️ Control Options:
|
|
|
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
Greptile SummaryThis PR normalizes mapping-form reasoning effort before Fireworks Chat requests are built
Confidence Score: 5/5The PR appears safe to merge because the mapping preserves existing scalar behavior and correctly flattens the structured input No actionable correctness, security, rules-compliance, or build failure remains in the changed code Important Files Changed
Reviews (1): Last reviewed commit: "refactor(fireworks_ai): extract reasonin..." | Re-trigger Greptile |
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 232233f. Configure here.
TLDR
Problem this solves:
reasoning_effortas{"effort": ..., "summary": ...}How it solves it:
_map_reasoning_efforthelper pullseffortout of dict-form valuesTrue-> medium,False-> none,auto-> drop rulesUser Flow
Before: a developer running Claude Code through a proxy that has reasoning summaries turned on gets a 400 on every prompt to a Fireworks reasoning model once extended thinking is on
fireworks_ai/accounts/fireworks/models/deepseek-v4p1-flashand setsLITELLM_REASONING_AUTO_SUMMARY=trueANTHROPIC_BASE_URL=https://litellm-domain,ANTHROPIC_MODEL=deepseek-v4p1-flash,MAX_THINKING_TOKENS=4096) and types a prompt"thinking": {"type": "enabled", "budget_tokens": 4096}Fireworks_aiException - Request body field 'reasoning_effort' is of type 'object', expected 'string', and Claude Code showsAPI Error: 400instead of a reply"reasoning_effort": {"effort": "medium", "summary": "detailed"}gets the same 400After: the same prompt gets a normal reply, with the reasoning kept
fireworks_ai/accounts/fireworks/models/deepseek-v4p1-flashand setsLITELLM_REASONING_AUTO_SUMMARY=trueANTHROPIC_BASE_URL=https://litellm-domain,ANTHROPIC_MODEL=deepseek-v4p1-flash,MAX_THINKING_TOKENS=4096) and types a prompt"thinking": {"type": "enabled", "budget_tokens": 4096}thinkingblock and the text, and Claude Code renders the replyreasoning_effortanswers 200 withreasoning_contenton the messageRelevant issues
Fixes #35649
Supersedes #35650, which reached the same fix first and gets credited here. #36363 fixed the Responses to Chat path only; this covers the Anthropic adapter path and a dict sent straight to chat completions
Affected release
Linear ticket
Resolves LIT-7843
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Shared setup: two proxies boot from the same config and env with two uvicorn workers each and no database, one from the merge base (878716f, port 55408) and one from the tip (232233f, port 25901). Every call below hits the real Fireworks API
Claude Code v2.1.273 is driven interactively under tmux:
ANTHROPIC_BASE_URL=http://127.0.0.1:$PORT ANTHROPIC_AUTH_TOKEN=$LITELLM_MASTER_KEY ANTHROPIC_MODEL=deepseek-v4p1-flash ANTHROPIC_DEFAULT_HAIKU_MODEL=deepseek-v4p1-flash MAX_THINKING_TOKENS=4096 claude, then the promptSay hello in one short sentenceBefore (878716f)
Claude Code on /v1/messages
API Error: 400 litellm.BadRequestError: Fireworks_aiException - {"error":{"message":"Request body field 'reasoning_effort' is of type 'object', expected 'string'","param":"reasoning_effort","code":"BAD_REQUEST","type":"error"}, ...}and no replycurl /v1/chat/completions with a dict reasoning_effort
curl http://127.0.0.1:55408/v1/chat/completions -H "Authorization: Bearer $LITELLM_MASTER_KEY" -H 'Content-Type: application/json' -d '{"model":"deepseek-v4p1-flash","messages":[{"role":"user","content":"Reply with the single word ok"}],"reasoning_effort":{"effort":"medium","summary":"detailed"},"max_tokens":300}'HTTP 400with{"error":{"message":"litellm.BadRequestError: Fireworks_aiException - {\"error\":{\"message\":\"Request body field 'reasoning_effort' is of type 'object', expected 'string'\",\"param\":\"reasoning_effort\",\"code\":\"BAD_REQUEST\",\"type\":\"error\"}, ...}. Received Model Group=deepseek-v4p1-flash\nAvailable Model Group Fallbacks=None"}}curl /v1/messages with thinking enabled
curl http://127.0.0.1:55408/v1/messages -H "x-api-key: $LITELLM_MASTER_KEY" -H 'anthropic-version: 2023-06-01' -H 'Content-Type: application/json' -d '{"model":"deepseek-v4p1-flash","max_tokens":1024,"thinking":{"type":"enabled","budget_tokens":1024},"messages":[{"role":"user","content":"Reply with the single word ok"}]}'HTTP 400with the sameRequest body field 'reasoning_effort' is of type 'object', expected 'string'error bodycurl /v1/responses with a reasoning object
curl http://127.0.0.1:55408/v1/responses -H "Authorization: Bearer $LITELLM_MASTER_KEY" -H 'Content-Type: application/json' -d '{"model":"deepseek-v4p1-flash","input":"Reply with the single word ok","reasoning":{"effort":"medium","summary":"detailed"},"max_output_tokens":300}'HTTP 200, output items["reasoning", "message"], message textok(already passing at the merge base: Fireworks has its own Responses route)After (232233f)
Claude Code on /v1/messages
Hello! I'm ready to help with your software engineering tasks.curl /v1/chat/completions with a dict reasoning_effort
HTTP 200,choices[0].message.contentisokandchoices[0].message.reasoning_contentis presentcurl /v1/messages with thinking enabled
HTTP 200, content blocks["thinking", "text"], textokcurl /v1/responses with a reasoning object
HTTP 200, output items["reasoning", "message"], message textokObservations:
reasoning_summaryon a Fireworks chat call 400s on both sides (extra input); this PR leaves it aloneType
🐛 Bug Fix
Caveats (if any)
Low
effortkey is dropped so Fireworks applies its model default, the same choice fix(fireworks-ai): handle dict-form reasoning_effort from Anthropic adapter #35650 made; forwarding it is the 400 this PR fixesthinking.summaryin the request; with summaries off the same /v1/messages call already passed at the merge baseUnable to access non-serverless model openai/gpt-oss-20bon every pipeline todaytest_proxy_budget_resettests fail at the merge base; main updated their assertions in a8fba14test_router_fallbacks_with_cooldowns_and_dynamic_credentials,test_async_fallbacks, andtest_generic_api_callback_sumologic_uses_ndjsonflap on unrelated pipelines and behave the same at the merge basetests/llm_translation/test_fireworks_ai_translation.pypassed in the same jobFinal Attestation
Link to Devin session: https://app.devin.ai/sessions/81933fbe62f649ae83f9b4aae0db02be
Open in Devin Desktop: https://app.devin.ai/desktop/session/81933fbe62f649ae83f9b4aae0db02be?variant=devin
Note
Low Risk
Scoped to Fireworks AI request param normalization for reasoning_effort; existing scalar/boolean/auto behavior is preserved via the shared helper.
Overview
Fixes Fireworks chat requests failing when
reasoning_effortarrives as an object (e.g.{"effort": "medium", "summary": "detailed"}from the Anthropic adapter) instead of the string Fireworks expects.FireworksAIConfig.map_openai_paramsnow routesreasoning_effortthrough_map_reasoning_effort, which reads the innereffortfrom dict-shaped values and keeps existing rules:True→"medium",False→"none","auto"omitted from the outbound request. Dicts with noeffortkey are dropped rather than forwarded.Tests cover dict flattening and the missing-
effortdrop behavior.Reviewed by Cursor Bugbot for commit 232233f. Bugbot is set up for automated code reviews on this repo. Configure here.