HTML Link Extractor
Extract anchor links from pasted HTML, resolve relative URLs with a base, filter same-origin or external links and export text, attributes and deduplicated URL lists.
Settings
Try the example or enter your own settings.
HTML is parsed in a detached inert template and never inserted into the visible page. Scripts are not executed, resources are not loaded, and URLs are not followed. Links are shown as text. Use a nonconfidential capture because third-party ad scripts can technically access inputs.
Using HTML Link Extractor
Extract anchor links from pasted HTML, resolve relative URLs with a base, filter same-origin or external links and export text, attributes and deduplicated URL lists.
- Paste static HTML source or open a UTF-8 file. The extractor reads a and area elements with href attributes.
- Set the original page URL to resolve relative links. Choose a filter and whether to keep duplicate destinations.
- Extract, inspect errors and rel tokens, then copy or download JSON, a resolved URL list or CSV rows.
Example
Try the included editable example. Set the original page URL to resolve relative links. Choose a filter and whether to keep duplicate destinations.
Questions & answers
Does it crawl a website?
No. You provide HTML text. The tool does not download URLs, visit destinations or check whether links work.
Can pasted scripts or images run?
The source is parsed in an unattached template. It is never moved into the page. Scripts remain inert and resources are not loaded by the extraction process. Output is rendered as plain text.
How are internal and external links defined?
Internal means the resolved URL has the same scheme, hostname and port as the supplied base URL. Without a base, absolute web URLs are classified as web rather than internal or external.
Does it honor an HTML base tag?
No. The tool notes its presence but uses only the base URL you enter. This makes relative URL resolution explicit.
What links are omitted?
Script-generated links, nested template contents, shadow DOM, image sources and CSS URLs are outside this extractor. Supported resolved schemes are HTTP, HTTPS, mailto and tel; credential URLs are excluded.
What limits apply?
Use 100,000 source characters and at most 2,000 anchor/area links. Link text is capped at 500 characters per entry. URL-list exports omit unresolved entries; JSON and the all-links table retain selected errors.
Help improve this tool
Report a problem or suggest an improvement
Describe the issue without pasting private tool input. Feedback goes to our admin inbox.
