Skip to content

fix(guardrails): don't add post_call output scan for MCP-only Presidio modes - #40571

Merged
yassin-berriai merged 3 commits into
mainfrom
litellm_presidio_mcp_mode_no_post_call_scan
Sep 16, 2026
Merged

yassin-berriai merged 3 commits into
mainfrom
litellm_presidio_mcp_mode_no_post_call_scan

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • MCP-only Presidio guardrails also scanned the final LLM answer
  • Blocked MCP tool query repeated in the answer turned a 200 into a 400
  • Tag-based mode objects were not recognized as MCP-only

How it solves it:

  • Modes made only of pre_mcp_call / during_mcp_call / post_mcp_call default to presidio_filter_scope: input
  • Works for a string, a list, and a tag-based mode (tags plus default)
  • Explicit presidio_filter_scope: both or output still scans the answer

User Flow

Before: a developer asks the model to run an MCP tool with a phone number in the query, and the whole request fails even though only the tool call should be blocked

  1. The proxy admin configures a Presidio guardrail with mode: pre_mcp_call and PHONE_NUMBER: BLOCK, plus one MCP server
  2. The developer sends POST http://localhost:4000/v1/responses with the MCP tool and the input "search for exactly this text: call me at 415-555-2671, repeat the exact query in your answer"
  3. The tool call is blocked as expected, and the model writes an answer that repeats the phone number
  4. The response comes back as HTTP 400 Blocked entity detected: PHONE_NUMBER by Guardrail: pre-mcp-pii-check, so the developer never sees the answer or the blocked tool result

After: the same request returns the answer, with the blocked tool result inside it

  1. The proxy admin configures the same guardrail and MCP server
  2. The developer sends the same POST http://localhost:4000/v1/responses
  3. The tool call is blocked as expected, and the model writes an answer that repeats the phone number
  4. The response comes back as HTTP 200 with the assistant message and a tool_execution_results item reading Tool call blocked: PII entity 'PHONE_NUMBER' detected by guardrail 'pre-mcp-pii-check'

Relevant issues

Pylon #8386

Affected release

Linear ticket

Resolves LIT-7459

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Shared setup: Presidio analyzer and anonymizer containers (mcr.microsoft.com/presidio-analyzer, mcr.microsoft.com/presidio-anonymizer), a local MCP server exposing one web_search tool on http://127.0.0.1:27459/mcp, and the proxy started with litellm --config config.yaml --port 20459 --detailed_debug. Real OpenAI calls to gpt-4.1-mini

model_list:
  - model_name: gpt-4.1-mini
    litellm_params:
      model: openai/gpt-4.1-mini
      api_key: os.environ/OPENAI_API_KEY

mcp_servers:
  demo:
    url: http://127.0.0.1:27459/mcp
    transport: http

guardrails:
  - guardrail_name: pre-mcp-pii-check
    litellm_params:
      guardrail: presidio
      mode: pre_mcp_call
      presidio_analyzer_api_base: http://127.0.0.1:25459
      presidio_anonymizer_api_base: http://127.0.0.1:26459
      pii_entities_config:
        PHONE_NUMBER: BLOCK
        US_SSN: BLOCK
      default_on: true
      # case 2 only: presidio_filter_scope: both

general_settings:
  master_key: sk-1234

Request used in every run (req.json):

{
  "model": "gpt-4.1-mini",
  "input": "Use the web_search tool to search for exactly this text: \"call me at 415-555-2671\". Do not modify the query. In your final answer, always repeat the exact query text you searched for, verbatim, even if the tool failed.",
  "tools": [{"type": "mcp", "server_label": "litellm", "server_url": "litellm_proxy", "require_approval": "never"}]
}

Before (29a712b)

MCP-only mode, default scope

  1. curl -s -i http://127.0.0.1:20459/v1/responses -H 'Authorization: Bearer sk-1234' -H 'content-type: application/json' -d @req.json
  2. Observed:
HTTP/1.1 400 Bad Request
{"error":{"message":"Blocked entity detected: PHONE_NUMBER by Guardrail: pre-mcp-pii-check. This entity is not allowed to be used in this request.","type":"invalid_request_error","param":null,"code":"400"}}
  1. Proxy log shows the tool call was blocked first, then the post-call scan of the assistant answer raised the 400:
async_post_call_success_hook response: ... text='I searched for the exact text: "call me at 415-555-2671" but the search was blocked ...'
... "result": "Tool call blocked: PII entity 'PHONE_NUMBER' detected by guardrail 'pre-mcp-pii-check'. ..."
LiteLLM Proxy:ERROR: common_request_processing.py - _handle_llm_api_exception(): Exception occured - Blocked entity detected: PHONE_NUMBER by Guardrail: pre-mcp-pii-check.

MCP-only mode with explicit presidio_filter_scope: both

  1. Same curl against the config with presidio_filter_scope: both
  2. Observed:
HTTP/1.1 400 Bad Request
{"error":{"message":"Blocked entity detected: PHONE_NUMBER by Guardrail: pre-mcp-pii-check. This entity is not allowed to be used in this request.","type":"invalid_request_error","param":null,"code":"400"}}

After (8a43fed)

MCP-only mode, default scope

  1. curl -s -i http://127.0.0.1:20459/v1/responses -H 'Authorization: Bearer sk-1234' -H 'content-type: application/json' -d @req.json
  2. Observed (trimmed to the relevant output items):
HTTP/1.1 200 OK
{
  "status": "completed",
  "output": [
    {
      "type": "message",
      "role": "assistant",
      "content": [{"type": "output_text", "text": "I attempted to search for the exact query \"call me at 415-555-2671\" but was unable to proceed because it contains a phone number, which is considered personal identifying information and is restricted by policy. The exact query text you asked me to search for is: \"call me at 415-555-2671\"."}]
    },
    {"type": "mcp_tools_fetched", "status": "completed", "role": "system", "content": [{"type": "output_text", "text": "[\n  \"name='demo-web_search' ..."}]},
    {
      "type": "tool_execution_results",
      "status": "completed",
      "role": "system",
      "content": [{"type": "output_text", "text": "[\n  {\n    \"tool_call_id\": \"call_llZNp30zpNrXohMNwVNrgFhV\",\n    \"result\": \"Tool call blocked: PII entity 'PHONE_NUMBER' detected by guardrail 'pre-mcp-pii-check'. Blocked entity detected: PHONE_NUMBER by Guardrail: pre-mcp-pii-check. This entity is not allowed to be used in this request.\",\n    \"name\": \"demo-web_search\"\n  }\n]"}]
    }
  ]
}

MCP-only mode with explicit presidio_filter_scope: both

  1. Same curl against the config with presidio_filter_scope: both
  2. Observed, unchanged from Before since the admin asked for output scanning:
HTTP/1.1 400 Bad Request
{"error":{"message":"Blocked entity detected: PHONE_NUMBER by Guardrail: pre-mcp-pii-check. This entity is not allowed to be used in this request.","type":"invalid_request_error","param":null,"code":"400"}}

Type

🐛 Bug Fix

Caveats (if any)

Medium

  • Deployments relying on the implicit output scan of an MCP-only guardrail must now set presidio_filter_scope: both to keep it

Low

  • Tag-based mode counts as MCP-only only when every tag value and the default are MCP hooks; a mix keeps the old both default
  • An empty tag-based mode (no tags, no default) also keeps the old both default

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Link to Devin session: https://app.devin.ai/sessions/f346adfada8f43468f6d29bcc05630bd
Open in Devin Desktop: https://app.devin.ai/desktop/session/f346adfada8f43468f6d29bcc05630bd?variant=devin
Requested by: @yassin-berriai


Note

Medium Risk
Changes the default Presidio filter scope for MCP-only guardrails; deployments that relied on implicit LLM output scanning must set presidio_filter_scope: both explicitly.

Overview
Fixes a regression where Presidio guardrails configured for MCP-only event hooks (pre_mcp_call, during_mcp_call, post_mcp_call) still registered a post_call output scan by default (presidio_filter_scope: both). That could turn an otherwise successful response into HTTP 400 when the model repeated PII from a blocked MCP tool call in its answer.

initialize_presidio now infers presidio_filter_scope: input when no explicit scope is set and every configured hook in mode is MCP-only (string, list, or tag-based Mode). Explicit both or output still enables answer scanning.

Adds helpers _configured_event_hooks and _is_mcp_only_mode, plus a parametrized test that exercises post_call behavior across MCP-only vs mixed modes and explicit filter scopes.

Reviewed by Cursor Bugbot for commit 8a43fed. Bugbot is set up for automated code reviews on this repo. Configure here.

@devin-ai-integration
devin-ai-integration Bot requested a review from a team September 10, 2026 10:19
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@codspeed

codspeed Bot commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_presidio_mcp_mode_no_post_call_scan (8a43fed) with main (f0474bb)

Open in CodSpeed

@greptile-apps

greptile-apps Bot commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR changes Presidio initialization so guardrails configured exclusively for MCP events default to input-only filtering, preventing an unrelated post-call scan of the final model response. It supports string, list, and tag-based modes while retaining output scanning when explicitly requested. The latest update replaces structural matching in _configured_event_hooks with explicit isinstance returns without changing the validated behavior

Confidence Score: 5/5

The PR appears safe to merge because the latest refactor preserves validated mode handling and both previous findings are resolved

Tag-based modes are flattened correctly, the regression test checks observable response redaction, and explicit output-filter scopes retain the existing scan behavior

Important Files Changed
Filename Overview
litellm/proxy/guardrails/guardrail_initializers.py Detects MCP-only Presidio modes across supported mode shapes and defaults those configurations to input-only filtering
tests/test_litellm/proxy/guardrails/test_init_guardrails.py Exercises observable post-call redaction behavior for MCP-only, mixed, empty, explicit-scope, and ordinary modes

Reviews (3): Last reviewed commit: "fix(guardrails): use explicit returns in..." | Re-trigger Greptile

Comment thread litellm/proxy/guardrails/guardrail_initializers.py Outdated
Comment thread tests/test_litellm/proxy/guardrails/test_init_guardrails.py Outdated
@codecov

codecov Bot commented Sep 10, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@yuneng-berri
yuneng-berri deleted the branch main September 13, 2026 04:51
@yuneng-berri yuneng-berri reopened this Sep 13, 2026
@mateo-berri
mateo-berri changed the base branch from litellm_internal_staging to main September 14, 2026 20:10
@devin-ai-integration
devin-ai-integration Bot force-pushed the litellm_presidio_mcp_mode_no_post_call_scan branch from 7dfda63 to 6bfa1af Compare September 15, 2026 21:02
Comment thread litellm/proxy/guardrails/guardrail_initializers.py Fixed
jesus-berri and others added 2 commits September 16, 2026 16:15
…o modes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…MCP hooks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration
devin-ai-integration Bot force-pushed the litellm_presidio_mcp_mode_no_post_call_scan branch from 6bfa1af to 7b3582a Compare September 16, 2026 16:31
@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you all sign our Contributor License Agreement before we can accept your contribution.
1 out of 2 committers have signed the CLA.

✅ jesus-berri
❌ yassin-berriai
You have signed the CLA already but the status is still pending? Let us recheck it.

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai please re-review: tag-based Mode is now classified as MCP-only, the test checks output-scan behavior, and CodeQL return finding is fixed

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

Fixed in 7b3582a: every match arm returns explicitly and the fallback arm calls assert_never, so there is no implicit None return

Comment thread litellm/proxy/guardrails/guardrail_initializers.py Fixed
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai please re-review: _configured_event_hooks now uses explicit isinstance returns, so the CodeQL mixed-return finding is gone

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@veria-ai please review 8a43fed: MCP-only Presidio modes now default to input-only scanning unless presidio_filter_scope asks for output

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 8a43fed. Configure here.

@yassin-berriai
yassin-berriai merged commit 8aebd4f into main Sep 16, 2026
98 of 99 checks passed
@yassin-berriai
yassin-berriai deleted the litellm_presidio_mcp_mode_no_post_call_scan branch September 16, 2026 17:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants