7 ms·
You are right. I should add a warning. The "Web URL" regular expression has the advantage of catching "example.com/foo?q=bar#baz", not just "https://example.com
by networked 3y ago
You are right. I should add a warning. The "Web URL" regular expression has the advantage of catching "example.com/foo?q=bar#baz", not just "https://example.com/foo?q=bar#baz https://example.com/foo?q=bar#baz", but Gruber published it in 2014 (https://gist.github.com/gruber/8891611/revisions https://gist.github.com/gruber/8891611/revisions). I have not updated it.
I recommend normally using the script without -w and filtering the output for HTTP(S) URLs. The default regular expression (https://gist.github.com/gruber/249502 https://gist.github.com/gruber/249502) does not rely on a list of TLDs.
Edit: Added a warning to the original comment. Thanks for prompting me to.
Here is a version without the "Web URL" regex.
#! /bin/sh
# shellcheck disable=SC1112
set -eu
# The URL regular expression by John Gruber.
# https://gist.github.com/gruber/249502
re='(?i)\b((?:[a-z][\w-]+:(?:/{1,3}|[a-z0-9%])|www\d{0,3}[.]|[a-z0-9.\-]+[.][a-z]{2,4}/)(?:[^\s()<>]+|\(([^\s()<>]+|(\([^\s()<>]+\)))*\))+(?:\(([^\s()<>]+|(\([^\s()<>]+\)))*\)|[^\s`!()\[\]{};:'"'"'".,<>«»“”‘’]))'
grep -oP "$re" "$@"