Skip to content

perf: defer fastapi and tiktoken BPE imports out of import litellm - #41585

Merged
yassin-berriai merged 5 commits into
mainfrom
litellm_lazy_fastapi_bpe_imports
Sep 18, 2026
Merged

yassin-berriai merged 5 commits into
mainfrom
litellm_lazy_fastapi_bpe_imports

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • import litellm (so every lite CLI command) loads fastapi and starlette
  • It also decodes the tiktoken cl100k BPE table at import time
  • A customer sees lite --help take up to 20 s on a 1 CPU VM

How it solves it:

  • token_counter fetches the default encoding through the existing lazy loader
  • custom_guardrail checks for a fastapi HTTPException through a small proxy-side helper, only when classifying an exception
  • gcs_bucket imports CommonProxyErrors inside the two functions that use it
  • Import graph shrinks by 78 modules on import litellm and during lite --help

User Flow

Before: an operator on a small VM waits several seconds for any lite command, even --help

  1. They run lite --help on a 1 CPU VM
  2. The terminal sits blank for many seconds before the help text prints
  3. lite models list and lite keys list pay the same startup delay on every call

After: the same commands start sooner, though startup is still dominated by the SDK import

  1. They run lite --help on the same VM
  2. The help text prints sooner because the web framework and the tokenizer table are no longer loaded
  3. lite models list and lite keys list see the same saving on every call

Relevant issues

Reported by a customer via Pylon #8739

Affected release

Linear ticket

Resolves LIT-7990

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Screenshots / Proof of Fix

Both arms run on the same many-core dev box (absolute numbers are far below the customer's 1 CPU VM) and share one .venv at /home/ubuntu/repos/litellm/.venv. Before is a sibling worktree at the merge base with PYTHONPATH=<tree>:<tree>/enterprise pointing at it, After is this checkout at the tip with the same PYTHONPATH shape. The proxy legs use real Postgres, litellm/proxy/dev_config.yaml, and a real Anthropic call. The served build is identified by the process cwd (readlink /proc/<pid>/cwd) and its PYTHONPATH from /proc/<pid>/environ, not by changing the payloads

Shared commands and request bodies ($TREE is the arm's checkout, $PY is /home/ubuntu/repos/litellm/.venv/bin/python):

modules:   LITELLM_LOCAL_MODEL_COST_MAP=True PYTHONPATH=$TREE $PY -c "import sys, litellm; print('modules:', len(sys.modules)); print(sorted(m for m in ('fastapi','starlette','litellm.litellm_core_utils.default_encoding') if m in sys.modules))"
import:    for i in 1 2 3 4 5; do LITELLM_LOCAL_MODEL_COST_MAP=True PYTHONPATH=$TREE /usr/bin/time -f %e $PY -c "import litellm"; done
cli help:  for i in 1 2 3 4 5; do PYTHONPATH=$TREE /usr/bin/time -f %e $PY -c "import sys; sys.argv=['lite','--help']; from litellm.proxy.client.cli import cli; cli()" >/dev/null; done
lite bin:  for i in 1 2 3; do PYTHONPATH=$TREE /usr/bin/time -f %e /home/ubuntu/repos/litellm/.venv/bin/lite --help >/dev/null; done; PYTHONPATH=$TREE /home/ubuntu/repos/litellm/.venv/bin/lite --help | head -3
cli mods:  PYTHONPATH=$TREE $PY -c "import sys; sys.argv=['lite','--help']; from litellm.proxy.client.cli import cli
           try: cli()
           except SystemExit: pass
           print('modules:', len(sys.modules), sorted(m for m in ('fastapi','starlette','litellm.litellm_core_utils.default_encoding') if m in sys.modules))" | tail -1
tokenizer: LITELLM_LOCAL_MODEL_COST_MAP=True PYTHONPATH=$TREE $PY -c "import sys, litellm; from litellm import token_counter; print('litellm.litellm_core_utils.default_encoding' in sys.modules); print(token_counter(model=None, text='Reply with exactly one word: pong')); print('litellm.litellm_core_utils.default_encoding' in sys.modules)"

req_tok_anthropic.json: {"model":"anthropic-haiku-4-5","messages":[{"role":"user","content":"hello world"}]}
req_tok_openai.json:    {"model":"gpt-4o-mini","messages":[{"role":"user","content":"hello world"}]}
req_chat.json:          {"model":"anthropic-haiku-4-5","messages":[{"role":"user","content":"Reply with exactly one word: pong"}],"max_tokens":10}

Before (8fc9c46)

Modules pulled in by import litellm

  1. modules command above with TREE=/home/ubuntu/repos/litellm_base
  2. Output: modules: 2408 then ['fastapi', 'litellm.litellm_core_utils.default_encoding', 'starlette']

import litellm wall clock, five runs

  1. import command above
  2. Output: 3.54, 3.98, 6.63, 3.85, 2.68

lite --help wall clock and module count, five runs

  1. cli help command above
  2. Output: 3.28, 2.59, 2.69, 2.62, 2.74
  3. cli mods command above prints modules: 2882 ['fastapi', 'litellm.litellm_core_utils.default_encoding', 'starlette']
  4. lite bin command above, the installed lite entry point the customer runs: 2.56, 2.66, 2.51, then prints Usage: lite [OPTIONS] COMMAND [ARGS]... and LiteLLM Proxy CLI - Manage your LiteLLM proxy server

Default tokenizer path in the SDK

  1. tokenizer command above
  2. Output: True, 7, True (the BPE table is already decoded before any token is counted)

Proxy boot, token counting and a real Anthropic call

  1. cd /home/ubuntu/repos/litellm_base && PYTHONPATH="$PWD:$PWD/enterprise" /home/ubuntu/repos/litellm/.venv/bin/python litellm/proxy/proxy_cli.py --config litellm/proxy/dev_config.yaml --detailed_debug --port 4001 --use_v2_migration_resolver
  2. readlink /proc/<pid>/cwd -> /home/ubuntu/repos/litellm_base, /proc/<pid>/environ has PYTHONPATH=/home/ubuntu/repos/litellm_base:/home/ubuntu/repos/litellm_base/enterprise, curl -s -o /dev/null -w '%{http_code}' http://localhost:4001/health/liveliness -> 200
  3. curl -s -w '\nHTTP %{http_code}\n' http://localhost:4001/utils/token_counter -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d @req_tok_anthropic.json
  4. Output: {"total_tokens":9,"request_model":"anthropic-haiku-4-5","model_used":"claude-haiku-4-5","tokenizer_type":"huggingface_tokenizer","original_response":null,"error":false,"error_message":null,"status_code":null} then HTTP 200
  5. Same with -d @req_tok_openai.json: {"total_tokens":9,"request_model":"gpt-4o-mini","model_used":"gpt-4o-mini","tokenizer_type":"openai_tokenizer",...} then HTTP 200
  6. curl -s -w '\nHTTP %{http_code}\n' http://localhost:4001/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d @req_chat.json
  7. Output: {"id":"chatcmpl-4b7e9045-153c-46aa-8c97-25f6d884db9e","created":1789722830,"model":"anthropic-haiku-4-5","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"pong","role":"assistant",...}}],"usage":{"completion_tokens":5,"prompt_tokens":15,"total_tokens":20,...}} then HTTP 200

After (b10944e)

Modules pulled in by import litellm

  1. modules command above with TREE=/home/ubuntu/repos/litellm
  2. Output: modules: 2330 then []

import litellm wall clock, five runs

  1. import command above
  2. Output: 2.19, 2.53, 2.13, 2.68, 2.11

lite --help wall clock and module count, five runs

  1. cli help command above
  2. Output: 2.39, 2.49, 2.36, 2.76, 2.96
  3. cli mods command above prints modules: 2804 []
  4. lite bin command above: 2.33, 2.38, 2.34, then prints the same Usage: lite [OPTIONS] COMMAND [ARGS]... and LiteLLM Proxy CLI - Manage your LiteLLM proxy server

Default tokenizer path in the SDK

  1. tokenizer command above
  2. Output: False, 7, True (same count, the BPE table is decoded on the first default-tokenizer call instead of at import)

Proxy boot, token counting and a real Anthropic call

  1. cd /home/ubuntu/repos/litellm && PYTHONPATH="$PWD:$PWD/enterprise" .venv/bin/python litellm/proxy/proxy_cli.py --config litellm/proxy/dev_config.yaml --detailed_debug --port 4000
  2. readlink /proc/<pid>/cwd -> /home/ubuntu/repos/litellm, /proc/<pid>/environ has PYTHONPATH=/home/ubuntu/repos/litellm:/home/ubuntu/repos/litellm/enterprise, curl -s -o /dev/null -w '%{http_code}' http://localhost:4000/health/liveliness -> 200
  3. curl -s -w '\nHTTP %{http_code}\n' http://localhost:4000/utils/token_counter -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d @req_tok_anthropic.json
  4. Output: {"total_tokens":9,"request_model":"anthropic-haiku-4-5","model_used":"claude-haiku-4-5","tokenizer_type":"huggingface_tokenizer","original_response":null,"error":false,"error_message":null,"status_code":null} then HTTP 200
  5. Same with -d @req_tok_openai.json: {"total_tokens":9,"request_model":"gpt-4o-mini","model_used":"gpt-4o-mini","tokenizer_type":"openai_tokenizer",...} then HTTP 200
  6. curl -s -w '\nHTTP %{http_code}\n' http://localhost:4000/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d @req_chat.json
  7. Output: {"id":"chatcmpl-44ea2bf8-83a5-4c34-9d0f-9bb104a83316","created":1789722589,"model":"anthropic-haiku-4-5","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"pong","role":"assistant",...}}],"usage":{"completion_tokens":5,"prompt_tokens":15,"total_tokens":20,...}} then HTTP 200

Admin UI logs page on the tip proxy

  1. Open http://localhost:4000/ui/?page=logs, sign in with admin and the master key
  2. The Request Logs table lists the call from step 6 above as Success, anthropic/claude-haiku-4-5, 20 tokens, $0.000040
  3. Click the row: the detail pane shows request id chatcmpl-44ea2bf8-83a5-4c34-9d0f-9bb104a83316, provider anthropic, 15 prompt plus 5 completion tokens, duration 0.568 s

Admin UI request log detail for the Anthropic call served by the tip proxy

Type

🐛 Bug Fix

Caveats (if any)

Medium

  • This is a bandage: lite still runs the whole SDK __init__ before parsing argv
  • The remaining ~2 s is the router, proxy types and provider modules; the depth fix is tracked on LIT-7990

Low

  • A proxy deployment still loads fastapi a moment later, so no behavior change there
  • Guardrail HTTPException classification now imports fastapi on first call instead of at import
  • The first default-tokenizer call now decodes the BPE table on the calling thread instead of at import; the proxy already routes token counting through asyncify, and the SDK paid the same cost at import before
  • Review bots on b10944e: Greptile 5/5 with last reviewed commit b10944e, Veria check run passed with "No security issues found", CodeQL passed with no inline findings. Bugbot reviewed the previous tip 37091c2 clean; the bugbot run for b10944e got "Bugbot is paused, on-demand spend limit reached", so there is no Bugbot verdict for the merge commit itself. The merge only brought in main, the PR diff is unchanged since 37091c2
  • Review audit against the litellm caveat taxonomy on this diff at b10944e
    • C3, C7: no public symbol removed, litellm.token_counter, CustomGuardrail._is_guardrail_intervention and GCSBucketLogger keep their signatures and return types; token_counter.default_encoding had no importers in the repo
    • O3: when fastapi cannot be imported the helper returns False, so the exception is reported as guardrail_failed_to_respond, same as before, never as a block
    • O4: every caller of the classification helper lives under litellm/proxy, so there is no SDK-only sibling to extend
    • F2, F3: no provider or endpoint surface changes, the diff only moves imports
    • J5: the deferred imports are module lookups after first use; no new blocking work is added to a hot async path beyond the one-time BPE decode noted above
    • T1 to T5, T13: the import boundary test runs in a fresh subprocess, the fastapi-missing test restores builtins.__import__ through monkeypatch, the default encoding test compares against tiktoken.get_encoding("cl100k_base") rather than a pinned count, and the GCS tests assert the raised ValueError rather than a call
    • H1 to H9: no docs, no new source comments, no em dashes, no Optional, no bare dict or Any, no new constants, no dashboard, Prisma or CI workflow changes
    • A, B, D, E, G, K, L, M, N, P, Q, R, S: not applicable, the diff has no state, concurrency, auth, budget, streaming, UI, secret, OAuth, tenancy, sub-call, deploy or pagination surface

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Link to Devin session: https://app.devin.ai/sessions/9b9a6fc67d3a4fc08d8131ea9bf3e037
Open in Devin Desktop: https://app.devin.ai/desktop/session/9b9a6fc67d3a4fc08d8131ea9bf3e037?variant=devin
Requested by: @yassin-berriai

This defers FastAPI, Starlette, and the cl100k BPE table until the paths that use them run

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration
devin-ai-integration Bot requested a review from a team September 17, 2026 09:01
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

CLAassistant commented Sep 17, 2026 •

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you all sign our Contributor License Agreement before we can accept your contribution.
0 out of 2 committers have signed the CLA.

❌ yassin-berriai
❌ devin-ai-integration[bot]
You have signed the CLA already but the status is still pending? Let us recheck it.

@codspeed

codspeed Bot commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_lazy_fastapi_bpe_imports (b10944e) with main (8fc9c46)

Open in CodSpeed

@greptile-apps

greptile-apps Bot commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge, with all previous findings resolved and no new actionable issues in the current changeset

Summary

This PR reduces LiteLLM SDK and CLI startup work by deferring FastAPI and default tiktoken encoding imports until first use

  • Moves FastAPI exception classification behind a lazily loaded proxy helper
  • Resolves the default tokenizer through the existing lazy loader
  • Defers proxy type imports in GCS logging paths
  • Adds regression coverage for import boundaries and preserved runtime behavior

Reviews (5) · Last reviewed commit: "Merge remote-tracking branch 'origin/mai..."

Comment thread litellm/integrations/custom_guardrail.py Outdated
Comment thread tests/test_litellm/litellm_core_utils/test_token_counter.py Outdated
Comment thread tests/test_litellm/test_lazy_imports.py
@codecov

codecov Bot commented Sep 17, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai please re-review at 37091c2, a clean merge of main on top of the previously scored commit

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@veria-ai please review 37091c2: deferred fastapi and tiktoken imports, guardrail HTTPException classification helper under proxy, GCS premium gating imports moved into functions

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 37091c2. Configure here.

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor

cursor Bot commented Sep 18, 2026

Copy link
Copy Markdown
Contributor

Bugbot is paused — on-demand spend limit reached

Bugbot uses usage-based billing for this team and has hit its on-demand spend limit.

A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue.

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@veria-ai please review b10944e: clean merge of main on top of the previously reviewed deferred import change

@yassin-berriai
yassin-berriai merged commit d42f448 into main Sep 18, 2026
97 checks passed
@yassin-berriai
yassin-berriai deleted the litellm_lazy_fastapi_bpe_imports branch September 18, 2026 18:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants