perf: defer fastapi and tiktoken BPE imports out of import litellm - #41585
Conversation
This defers FastAPI, Starlette, and the cl100k BPE table until the paths that use them run Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
🤖 Devin AI EngineerI'll be helping with this pull request! Here's what you should know: ✅ I will automatically:
Note: I can only respond to comments from users who have write access to this repository. ⚙️ Control Options:
|
|
|
|
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
bugbot run |
|
bugbot run |
|
@greptileai please re-review at 37091c2, a clean merge of main on top of the previously scored commit |
|
@veria-ai please review 37091c2: deferred fastapi and tiktoken imports, guardrail HTTPException classification helper under proxy, GCS premium gating imports moved into functions |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 37091c2. Configure here.
|
bugbot run |
Bugbot is paused — on-demand spend limit reachedBugbot uses usage-based billing for this team and has hit its on-demand spend limit. A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue. |
|
@veria-ai please review b10944e: clean merge of main on top of the previously reviewed deferred import change |
TLDR
Problem this solves:
import litellm(so everyliteCLI command) loads fastapi and starlettelite --helptake up to 20 s on a 1 CPU VMHow it solves it:
token_counterfetches the default encoding through the existing lazy loadercustom_guardrailchecks for a fastapiHTTPExceptionthrough a small proxy-side helper, only when classifying an exceptiongcs_bucketimportsCommonProxyErrorsinside the two functions that use itimport litellmand duringlite --helpUser Flow
Before: an operator on a small VM waits several seconds for any
litecommand, even--helplite --helpon a 1 CPU VMlite models listandlite keys listpay the same startup delay on every callAfter: the same commands start sooner, though startup is still dominated by the SDK import
lite --helpon the same VMlite models listandlite keys listsee the same saving on every callRelevant issues
Reported by a customer via Pylon #8739
Affected release
Linear ticket
Resolves LIT-7990
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Screenshots / Proof of Fix
Both arms run on the same many-core dev box (absolute numbers are far below the customer's 1 CPU VM) and share one
.venvat/home/ubuntu/repos/litellm/.venv. Before is a sibling worktree at the merge base withPYTHONPATH=<tree>:<tree>/enterprisepointing at it, After is this checkout at the tip with the samePYTHONPATHshape. The proxy legs use real Postgres,litellm/proxy/dev_config.yaml, and a real Anthropic call. The served build is identified by the process cwd (readlink /proc/<pid>/cwd) and itsPYTHONPATHfrom/proc/<pid>/environ, not by changing the payloadsShared commands and request bodies (
$TREEis the arm's checkout,$PYis/home/ubuntu/repos/litellm/.venv/bin/python):Before (8fc9c46)
Modules pulled in by
import litellmmodulescommand above withTREE=/home/ubuntu/repos/litellm_basemodules: 2408then['fastapi', 'litellm.litellm_core_utils.default_encoding', 'starlette']import litellmwall clock, five runsimportcommand above3.54,3.98,6.63,3.85,2.68lite --helpwall clock and module count, five runscli helpcommand above3.28,2.59,2.69,2.62,2.74cli modscommand above printsmodules: 2882 ['fastapi', 'litellm.litellm_core_utils.default_encoding', 'starlette']lite bincommand above, the installedliteentry point the customer runs:2.56,2.66,2.51, then printsUsage: lite [OPTIONS] COMMAND [ARGS]...andLiteLLM Proxy CLI - Manage your LiteLLM proxy serverDefault tokenizer path in the SDK
tokenizercommand aboveTrue,7,True(the BPE table is already decoded before any token is counted)Proxy boot, token counting and a real Anthropic call
cd /home/ubuntu/repos/litellm_base && PYTHONPATH="$PWD:$PWD/enterprise" /home/ubuntu/repos/litellm/.venv/bin/python litellm/proxy/proxy_cli.py --config litellm/proxy/dev_config.yaml --detailed_debug --port 4001 --use_v2_migration_resolverreadlink /proc/<pid>/cwd->/home/ubuntu/repos/litellm_base,/proc/<pid>/environhasPYTHONPATH=/home/ubuntu/repos/litellm_base:/home/ubuntu/repos/litellm_base/enterprise,curl -s -o /dev/null -w '%{http_code}' http://localhost:4001/health/liveliness->200curl -s -w '\nHTTP %{http_code}\n' http://localhost:4001/utils/token_counter -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d @req_tok_anthropic.json{"total_tokens":9,"request_model":"anthropic-haiku-4-5","model_used":"claude-haiku-4-5","tokenizer_type":"huggingface_tokenizer","original_response":null,"error":false,"error_message":null,"status_code":null}thenHTTP 200-d @req_tok_openai.json:{"total_tokens":9,"request_model":"gpt-4o-mini","model_used":"gpt-4o-mini","tokenizer_type":"openai_tokenizer",...}thenHTTP 200curl -s -w '\nHTTP %{http_code}\n' http://localhost:4001/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d @req_chat.json{"id":"chatcmpl-4b7e9045-153c-46aa-8c97-25f6d884db9e","created":1789722830,"model":"anthropic-haiku-4-5","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"pong","role":"assistant",...}}],"usage":{"completion_tokens":5,"prompt_tokens":15,"total_tokens":20,...}}thenHTTP 200After (b10944e)
Modules pulled in by
import litellmmodulescommand above withTREE=/home/ubuntu/repos/litellmmodules: 2330then[]import litellmwall clock, five runsimportcommand above2.19,2.53,2.13,2.68,2.11lite --helpwall clock and module count, five runscli helpcommand above2.39,2.49,2.36,2.76,2.96cli modscommand above printsmodules: 2804 []lite bincommand above:2.33,2.38,2.34, then prints the sameUsage: lite [OPTIONS] COMMAND [ARGS]...andLiteLLM Proxy CLI - Manage your LiteLLM proxy serverDefault tokenizer path in the SDK
tokenizercommand aboveFalse,7,True(same count, the BPE table is decoded on the first default-tokenizer call instead of at import)Proxy boot, token counting and a real Anthropic call
cd /home/ubuntu/repos/litellm && PYTHONPATH="$PWD:$PWD/enterprise" .venv/bin/python litellm/proxy/proxy_cli.py --config litellm/proxy/dev_config.yaml --detailed_debug --port 4000readlink /proc/<pid>/cwd->/home/ubuntu/repos/litellm,/proc/<pid>/environhasPYTHONPATH=/home/ubuntu/repos/litellm:/home/ubuntu/repos/litellm/enterprise,curl -s -o /dev/null -w '%{http_code}' http://localhost:4000/health/liveliness->200curl -s -w '\nHTTP %{http_code}\n' http://localhost:4000/utils/token_counter -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d @req_tok_anthropic.json{"total_tokens":9,"request_model":"anthropic-haiku-4-5","model_used":"claude-haiku-4-5","tokenizer_type":"huggingface_tokenizer","original_response":null,"error":false,"error_message":null,"status_code":null}thenHTTP 200-d @req_tok_openai.json:{"total_tokens":9,"request_model":"gpt-4o-mini","model_used":"gpt-4o-mini","tokenizer_type":"openai_tokenizer",...}thenHTTP 200curl -s -w '\nHTTP %{http_code}\n' http://localhost:4000/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d @req_chat.json{"id":"chatcmpl-44ea2bf8-83a5-4c34-9d0f-9bb104a83316","created":1789722589,"model":"anthropic-haiku-4-5","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"pong","role":"assistant",...}}],"usage":{"completion_tokens":5,"prompt_tokens":15,"total_tokens":20,...}}thenHTTP 200Admin UI logs page on the tip proxy
adminand the master keyanthropic/claude-haiku-4-5, 20 tokens, $0.000040chatcmpl-44ea2bf8-83a5-4c34-9d0f-9bb104a83316, provider anthropic, 15 prompt plus 5 completion tokens, duration 0.568 sType
🐛 Bug Fix
Caveats (if any)
Medium
litestill runs the whole SDK__init__before parsing argvLow
HTTPExceptionclassification now imports fastapi on first call instead of at importasyncify, and the SDK paid the same cost at import beforebugbot runfor b10944e got "Bugbot is paused, on-demand spend limit reached", so there is no Bugbot verdict for the merge commit itself. The merge only brought in main, the PR diff is unchanged since 37091c2litellm.token_counter,CustomGuardrail._is_guardrail_interventionandGCSBucketLoggerkeep their signatures and return types;token_counter.default_encodinghad no importers in the repoguardrail_failed_to_respond, same as before, never as a blocklitellm/proxy, so there is no SDK-only sibling to extendbuiltins.__import__through monkeypatch, the default encoding test compares againsttiktoken.get_encoding("cl100k_base")rather than a pinned count, and the GCS tests assert the raisedValueErrorrather than a callOptional, no baredictorAny, no new constants, no dashboard, Prisma or CI workflow changesFinal Attestation
Link to Devin session: https://app.devin.ai/sessions/9b9a6fc67d3a4fc08d8131ea9bf3e037
Open in Devin Desktop: https://app.devin.ai/desktop/session/9b9a6fc67d3a4fc08d8131ea9bf3e037?variant=devin
Requested by: @yassin-berriai