Skip to content

fix(fireworks_ai): flatten dict-form reasoning_effort to its effort string - #41335

Merged
mateo-berri merged 2 commits into
mainfrom
litellm_fireworks_dict_reasoning_effort
Sep 16, 2026
Merged

mateo-berri merged 2 commits into
mainfrom
litellm_fireworks_dict_reasoning_effort

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Anthropic adapter emits reasoning_effort as {"effort": ..., "summary": ...}
  • Fireworks transformer forwards that dict unchanged, Fireworks rejects the object

How it solves it:

  • New _map_reasoning_effort helper pulls effort out of dict-form values
  • Same helper keeps the existing True -> medium, False -> none, auto -> drop rules

User Flow

Before: a developer running Claude Code through a proxy that has reasoning summaries turned on gets a 400 on every prompt to a Fireworks reasoning model once extended thinking is on

  1. The proxy admin adds a Fireworks reasoning model such as fireworks_ai/accounts/fireworks/models/deepseek-v4p1-flash and sets LITELLM_REASONING_AUTO_SUMMARY=true
  2. The developer points Claude Code at the proxy (ANTHROPIC_BASE_URL=https://litellm-domain, ANTHROPIC_MODEL=deepseek-v4p1-flash, MAX_THINKING_TOKENS=4096) and types a prompt
  3. Claude Code sends POST https://litellm-domain/v1/messages with "thinking": {"type": "enabled", "budget_tokens": 4096}
  4. The proxy answers 400 Fireworks_aiException - Request body field 'reasoning_effort' is of type 'object', expected 'string', and Claude Code shows API Error: 400 instead of a reply
  5. A direct POST https://litellm-domain/v1/chat/completions carrying "reasoning_effort": {"effort": "medium", "summary": "detailed"} gets the same 400

After: the same prompt gets a normal reply, with the reasoning kept

  1. The proxy admin adds a Fireworks reasoning model such as fireworks_ai/accounts/fireworks/models/deepseek-v4p1-flash and sets LITELLM_REASONING_AUTO_SUMMARY=true
  2. The developer points Claude Code at the proxy (ANTHROPIC_BASE_URL=https://litellm-domain, ANTHROPIC_MODEL=deepseek-v4p1-flash, MAX_THINKING_TOKENS=4096) and types a prompt
  3. Claude Code sends POST https://litellm-domain/v1/messages with "thinking": {"type": "enabled", "budget_tokens": 4096}
  4. The proxy answers 200 with a thinking block and the text, and Claude Code renders the reply
  5. The direct POST https://litellm-domain/v1/chat/completions with the same dict-form reasoning_effort answers 200 with reasoning_content on the message

Relevant issues

Fixes #35649

Supersedes #35650, which reached the same fix first and gets credited here. #36363 fixed the Responses to Chat path only; this covers the Anthropic adapter path and a dict sent straight to chat completions

Affected release

Linear ticket

Resolves LIT-7843

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Shared setup: two proxies boot from the same config and env with two uvicorn workers each and no database, one from the merge base (878716f, port 55408) and one from the tip (232233f, port 25901). Every call below hits the real Fireworks API

model_list:
  - model_name: deepseek-v4p1-flash
    litellm_params:
      model: fireworks_ai/accounts/fireworks/models/deepseek-v4p1-flash
      api_key: os.environ/FIREWORKS_AI_API_KEY
general_settings:
  master_key: os.environ/LITELLM_MASTER_KEY
export LITELLM_REASONING_AUTO_SUMMARY=true LITELLM_LOCAL_MODEL_COST_MAP=True
python litellm/proxy/proxy_cli.py --config qa_config.yaml --port $PORT --num_workers 2 --detailed_debug

Claude Code v2.1.273 is driven interactively under tmux: ANTHROPIC_BASE_URL=http://127.0.0.1:$PORT ANTHROPIC_AUTH_TOKEN=$LITELLM_MASTER_KEY ANTHROPIC_MODEL=deepseek-v4p1-flash ANTHROPIC_DEFAULT_HAIKU_MODEL=deepseek-v4p1-flash MAX_THINKING_TOKENS=4096 claude, then the prompt Say hello in one short sentence

Before (878716f)

Claude Code on /v1/messages

  1. Launch Claude Code against port 55408 and send the prompt
  2. Claude Code shows API Error: 400 litellm.BadRequestError: Fireworks_aiException - {"error":{"message":"Request body field 'reasoning_effort' is of type 'object', expected 'string'","param":"reasoning_effort","code":"BAD_REQUEST","type":"error"}, ...} and no reply

pr41335-232233f654-lit7843-before-claude-code.png

curl /v1/chat/completions with a dict reasoning_effort

  1. curl http://127.0.0.1:55408/v1/chat/completions -H "Authorization: Bearer $LITELLM_MASTER_KEY" -H 'Content-Type: application/json' -d '{"model":"deepseek-v4p1-flash","messages":[{"role":"user","content":"Reply with the single word ok"}],"reasoning_effort":{"effort":"medium","summary":"detailed"},"max_tokens":300}'
  2. HTTP 400 with {"error":{"message":"litellm.BadRequestError: Fireworks_aiException - {\"error\":{\"message\":\"Request body field 'reasoning_effort' is of type 'object', expected 'string'\",\"param\":\"reasoning_effort\",\"code\":\"BAD_REQUEST\",\"type\":\"error\"}, ...}. Received Model Group=deepseek-v4p1-flash\nAvailable Model Group Fallbacks=None"}}

curl /v1/messages with thinking enabled

  1. curl http://127.0.0.1:55408/v1/messages -H "x-api-key: $LITELLM_MASTER_KEY" -H 'anthropic-version: 2023-06-01' -H 'Content-Type: application/json' -d '{"model":"deepseek-v4p1-flash","max_tokens":1024,"thinking":{"type":"enabled","budget_tokens":1024},"messages":[{"role":"user","content":"Reply with the single word ok"}]}'
  2. HTTP 400 with the same Request body field 'reasoning_effort' is of type 'object', expected 'string' error body

curl /v1/responses with a reasoning object

  1. curl http://127.0.0.1:55408/v1/responses -H "Authorization: Bearer $LITELLM_MASTER_KEY" -H 'Content-Type: application/json' -d '{"model":"deepseek-v4p1-flash","input":"Reply with the single word ok","reasoning":{"effort":"medium","summary":"detailed"},"max_output_tokens":300}'
  2. HTTP 200, output items ["reasoning", "message"], message text ok (already passing at the merge base: Fireworks has its own Responses route)

After (232233f)

Claude Code on /v1/messages

  1. Launch Claude Code against port 25901 and send the same prompt
  2. Claude Code renders Hello! I'm ready to help with your software engineering tasks.

pr41335-232233f654-lit7843-after-claude-code.png

curl /v1/chat/completions with a dict reasoning_effort

  1. The same curl against port 25901
  2. HTTP 200, choices[0].message.content is ok and choices[0].message.reasoning_content is present

curl /v1/messages with thinking enabled

  1. The same curl against port 25901
  2. HTTP 200, content blocks ["thinking", "text"], text ok

curl /v1/responses with a reasoning object

  1. The same curl against port 25901
  2. HTTP 200, output items ["reasoning", "message"], message text ok

Observations:

  • /v1/responses passed at the merge base too; this PR leaves it alone
  • With summaries off, /v1/messages with thinking passes on both sides
  • reasoning_summary on a Fireworks chat call 400s on both sides (extra input); this PR leaves it alone

Type

🐛 Bug Fix

Caveats (if any)

Low

  • A dict without an effort key is dropped so Fireworks applies its model default, the same choice fix(fireworks-ai): handle dict-form reasoning_effort from Anthropic adapter #35650 made; forwarding it is the 400 this PR fixes
  • Only fires with reasoning summaries on or a thinking.summary in the request; with summaries off the same /v1/messages call already passed at the merge base
  • CircleCI at the tip is red only on tests this PR does not touch, each red fleet-wide today or already at the merge base
    • 22 Together AI tests fail with Unable to access non-serverless model openai/gpt-oss-20b on every pipeline today
    • 4 test_proxy_budget_reset tests fail at the merge base; main updated their assertions in a8fba14
    • test_router_fallbacks_with_cooldowns_and_dynamic_credentials, test_async_fallbacks, and test_generic_api_callback_sumologic_uses_ndjson flap on unrelated pipelines and behave the same at the merge base
    • tests/llm_translation/test_fireworks_ai_translation.py passed in the same job

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Link to Devin session: https://app.devin.ai/sessions/81933fbe62f649ae83f9b4aae0db02be
Open in Devin Desktop: https://app.devin.ai/desktop/session/81933fbe62f649ae83f9b4aae0db02be?variant=devin


Note

Low Risk
Scoped to Fireworks AI request param normalization for reasoning_effort; existing scalar/boolean/auto behavior is preserved via the shared helper.

Overview
Fixes Fireworks chat requests failing when reasoning_effort arrives as an object (e.g. {"effort": "medium", "summary": "detailed"} from the Anthropic adapter) instead of the string Fireworks expects.

FireworksAIConfig.map_openai_params now routes reasoning_effort through _map_reasoning_effort, which reads the inner effort from dict-shaped values and keeps existing rules: True → "medium", False → "none", "auto" omitted from the outbound request. Dicts with no effort key are dropped rather than forwarded.

Tests cover dict flattening and the missing-effort drop behavior.

Reviewed by Cursor Bugbot for commit 232233f. Bugbot is set up for automated code reviews on this repo. Configure here.

…tring

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration
devin-ai-integration Bot requested a review from a team September 16, 2026 00:40
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@codspeed

codspeed Bot commented Sep 16, 2026

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_fireworks_dict_reasoning_effort (232233f) with main (878716f)

Open in CodSpeed

@codecov

codecov Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@greptile-apps

greptile-apps Bot commented Sep 16, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR normalizes mapping-form reasoning effort before Fireworks Chat requests are built

  • Extracts the effort field while preserving existing boolean and auto handling
  • Adds regression coverage for mapping-form values and missing effort fields

Confidence Score: 5/5

The PR appears safe to merge because the mapping preserves existing scalar behavior and correctly flattens the structured input

No actionable correctness, security, rules-compliance, or build failure remains in the changed code

Important Files Changed
Filename Overview
litellm/llms/fireworks_ai/chat/transformation.py Adds focused normalization that converts structured reasoning effort into Fireworks' flat Chat parameter
tests/test_litellm/llms/fireworks_ai/chat/test_fireworks_ai_chat_transformation.py Adds focused regression tests for valid and incomplete mapping-form reasoning effort

Reviews (1): Last reviewed commit: "refactor(fireworks_ai): extract reasonin..." | Re-trigger Greptile

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 232233f. Configure here.

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@mateo-berri
mateo-berri merged commit 737929f into main Sep 16, 2026
124 of 141 checks passed
@mateo-berri
mateo-berri deleted the litellm_fireworks_dict_reasoning_effort branch September 16, 2026 20:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Fireworks AI transformer passes dict-form reasoning_effort to API, causing 400

2 participants