Skip to content

Preserve string contents during classic prompt removal - #15386

Draft
tlikhit wants to merge 1 commit into
ipython:mainfrom
tlikhit:fix/preserve-string-prompts
Draft

tlikhit wants to merge 1 commit into
ipython:mainfrom
tlikhit:fix/preserve-string-prompts

Conversation

@tlikhit

@tlikhit tlikhit commented Sep 7, 2026 •

Copy link
Copy Markdown

Summary

Preserve literal >>> and ... text inside Python strings while still removing the outer prompts from pasted Python sessions.

Fixes #14600. The original bare-string examples now work, but assigned multiline strings still lose content in IPython 9.17.1 and main at e18b391ad.

Problem

Before Python executes a cell, IPython transforms its input. One step removes prompts such as >>> and ..., so code copied from an interactive session can run.

Those characters can also be part of a string. The current code uses a regular expression to find ''' and """, then tracks whether subsequent lines are inside a string. This approximate quote-tracking logic is the "triple-quote regex heuristic."

One of its protection rules requires the text before the opening quotes, after removing any prompt and surrounding whitespace, to be empty or only a string prefix such as f or r. An assignment such as text = f""" fails that check, so literal prompt text inside the string can be stripped. For example, this is ordinary Python code, not a pasted session:

text = f"""
>>> {1}
... tail
"""
Value of text
Expected "\n>>> 1\n... tail\n"
Before this fix "\n1\ntail\n"

The cell runs without a syntax error, but its string value changes silently.

Fix

Use Python's tokenizer to locate strings instead of the triple-quote heuristic. The opening line determines how to handle each string:

  • No prompt on the opening line: preserve its original contents, including any literal prompt text.
  • Prompt on the opening line: remove one outer prompt layer from the pasted code, leaving any inner prompt text intact.

For example, this pasted session must still have its outer prompts removed:

>>> text = """
... >>> literal
... ... tail
... """

The resulting string is "\n>>> literal\n... tail\n".

The tokenizer reads a temporary copy with prompts removed. The transformer retains the original lines for output. It also protects incomplete strings and resumes tokenization after lexical errors in shell or magic input.

Related behavior changes

Keep statements inside their enclosing block

Previously, indentation cleanup treated the code before and after a protected string as separate chunks. Consider a function and docstring followed by a pasted return line:

def greet():
    """A greeting function."""
>>>     return "hello"

After removing >>> , the old code also removed the four spaces before return, moving it outside the function. The fix uses the whole cell's indentation context, so the result is:

def greet():
    """A greeting function."""
    return "hello"

Remove outer prompts from a bare string

A bare string is a string expression without an assignment such as text =. For this pasted session:

>>> """
... >>> example text
... """

The old code left the entire input untouched, including the outer prompts. The fix removes those prompts but retains the literal >>> inside the string:

"""
>>> example text
"""

The existing CLASSIC_PROMPT_L3 test expected the old, invalid transformed source. Its expected result is updated to reflect this correction.

Documentation and scope

A release note documents the changes. IPython's numbered In [n]: prompts use a separate removal path, which this PR does not change.

Test coverage

Add parameterized regression tests and controls for:

  • Both triple-quote delimiters, string prefixes, and assignment/expression contexts.
  • Plain Python, doctest/xdoctest pastes, and mixed prompted/unprompted input.
  • Escaped quotes, misleading delimiters in comments, and incomplete strings.
  • Block indentation, LF/CRLF line endings, and shell/magic input.
  • Nested f-strings, Python 3.14 template strings, and InteractiveShell.run_cell.

Valid examples check both transformed source and resulting values, not just whether the code parses.

Review focus

This remains a draft for feedback on the opening-line rule and the changed bare-string expectation. Tokenization adds work when a cell contains a line matching >>>; cells without such a line return early.

AI-assisted implementation: 🤖 🤖

@tlikhit
tlikhit force-pushed the fix/preserve-string-prompts branch from 406efb8 to 6ede4c7 Compare September 7, 2026 23:58

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Cell Transformation Issue for Multline Strings containing prompt strings

1 participant