Resolve HTML links and images against the document base URL - #2501
Open
Kunpeng Xie (pentaoa) wants to merge 1 commit into
Open
Kunpeng Xie (pentaoa) wants to merge 1 commit into
Kunpeng Xie (pentaoa) wants to merge 1 commit into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #2500.
A page at
https://example.com/docs/page.htmlcontaining<a href="guide.html">currently produces a relative Markdown link. Its target is lost when the output is used outside the original page. Resolve rendered links and images against the supplied source URL, honoring the first<base href>when present.The change is confined to
HtmlConverterand resolvesa[href],img[src], andimg[data-src]before Markdown conversion. Existing link filtering and data-URI handling remain in the Markdown converter. Relative URLs stay relative when there is no source/base information; a malformed URL does not discard the document.Tests cover source and explicit bases, relative bases, first-base precedence, normal/lazy images, encoded paths, absolute links, malformed bases, and
MarkItDown.convert_response()using the final response URL. Six of the initial seven regression cases failed before the fix. The final HTML and RSS suites pass 142 tests. Black andgit diff --checkpass; no dependencies or network access were added. The full multi-format suite was not run.AI assistance: OpenAI Codex was used for investigation, implementation, and local test execution.