4 ms·
urls This script extracts URLs from a text input stream or text files using John Gruber's regular expressions. It requires GNU grep. If your system's default g
by networked 3y ago
urls
This script extracts URLs from a text input stream or text files using John Gruber's regular expressions. It requires GNU grep. If your system's default grep command isn't GNU, install ggrep and modify the script accordingly.
Save the script, make it executable, and try
$ ./urls urls
The output should be
https://gist.github.com/gruber/249502
https://gist.github.com/gruber/8891611
Usage:
urls [-w] [<grep arg> ...]
Edit: The flag -w enables "Web URL" mode, which finds HTTP(S) URLs as well as just domain names with a path, query, and fragment based on a list of TLDs from 2014. Warning: it will miss new TLDs. You can update the list from https://data.iana.org/TLD/tlds-alpha-by-domain.txt https://data.iana.org/TLD/tlds-alpha-by-domain.txt.
Source code:
#! /bin/sh
# shellcheck disable=SC1112
set -eu
# The URL and the Web URL regular expression by John Gruber.
# https://gist.github.com/gruber/249502
re_all='(?i)\b((?:[a-z][\w-]+:(?:/{1,3}|[a-z0-9%])|www\d{0,3}[.]|[a-z0-9.\-]+[.][a-z]{2,4}/)(?:[^\s()<>]+|\(([^\s()<>]+|(\([^\s()<>]+\)))*\))+(?:\(([^\s()<>]+|(\([^\s()<>]+\)))*\)|[^\s`!()\[\]{};:'"'"'".,<>«»“”‘’]))'
# https://gist.github.com/gruber/8891611
re_web='(?i)\b((?:https?:(?:/{1,3}|[a-z0-9%])|[a-z0-9.\-]+[.](?:com|net|org|edu|gov|mil|aero|asia|biz|cat|coop|info|int|jobs|mobi|museum|name|post|pro|tel|travel|xxx|ac|ad|ae|af|ag|ai|al|am|an|ao|aq|ar|as|at|au|aw|ax|az|ba|bb|bd|be|bf|bg|bh|bi|bj|bm|bn|bo|br|bs|bt|bv|bw|by|bz|ca|cc|cd|cf|cg|ch|ci|ck|cl|cm|cn|co|cr|cs|cu|cv|cx|cy|cz|dd|de|dj|dk|dm|do|dz|ec|ee|eg|eh|er|es|et|eu|fi|fj|fk|fm|fo|fr|ga|gb|gd|ge|gf|gg|gh|gi|gl|gm|gn|gp|gq|gr|gs|gt|gu|gw|gy|hk|hm|hn|hr|ht|hu|id|ie|il|im|in|io|iq|ir|is|it|je|jm|jo|jp|ke|kg|kh|ki|km|kn|kp|kr|kw|ky|kz|la|lb|lc|li|lk|lr|ls|lt|lu|lv|ly|ma|mc|md|me|mg|mh|mk|ml|mm|mn|mo|mp|mq|mr|ms|mt|mu|mv|mw|mx|my|mz|na|nc|ne|nf|ng|ni|nl|no|np|nr|nu|nz|om|pa|pe|pf|pg|ph|pk|pl|pm|pn|pr|ps|pt|pw|py|qa|re|ro|rs|ru|rw|sa|sb|sc|sd|se|sg|sh|si|sj|Ja|sk|sl|sm|sn|so|sr|ss|st|su|sv|sx|sy|sz|tc|td|tf|tg|th|tj|tk|tl|tm|tn|to|tp|tr|tt|tv|tw|tz|ua|ug|uk|us|uy|uz|va|vc|ve|vg|vi|vn|vu|wf|ws|ye|yt|yu|za|zm|zw)/)(?:[^\s()<>{}\[\]]+|\([^\s()]*?\([^\s()]+\)[^\s()]*?\)|\([^\s]+?\))+(?:\([^\s()]*?\([^\s()]+\)[^\s()]*?\)|\([^\s]+?\)|[^\s`!()\[\]{};:'"'"'".,<>?«»“”‘’])|(?:(?<!@)[a-z0-9]+(?:[.\-][a-z0-9]+)*[.](?:com|net|org|edu|gov|mil|aero|asia|biz|cat|coop|info|int|jobs|mobi|museum|name|post|pro|tel|travel|xxx|ac|ad|ae|af|ag|ai|al|am|an|ao|aq|ar|as|at|au|aw|ax|az|ba|bb|bd|be|bf|bg|bh|bi|bj|bm|bn|bo|br|bs|bt|bv|bw|by|bz|ca|cc|cd|cf|cg|ch|ci|ck|cl|cm|cn|co|cr|cs|cu|cv|cx|cy|cz|dd|de|dj|dk|dm|do|dz|ec|ee|eg|eh|er|es|et|eu|fi|fj|fk|fm|fo|fr|ga|gb|gd|ge|gf|gg|gh|gi|gl|gm|gn|gp|gq|gr|gs|gt|gu|gw|gy|hk|hm|hn|hr|ht|hu|id|ie|il|im|in|io|iq|ir|is|it|je|jm|jo|jp|ke|kg|kh|ki|km|kn|kp|kr|kw|ky|kz|la|lb|lc|li|lk|lr|ls|lt|lu|lv|ly|ma|mc|md|me|mg|mh|mk|ml|mm|mn|mo|mp|mq|mr|ms|mt|mu|mv|mw|mx|my|mz|na|nc|ne|nf|ng|ni|nl|no|np|nr|nu|nz|om|pa|pe|pf|pg|ph|pk|pl|pm|pn|pr|ps|pt|pw|py|qa|re|ro|rs|ru|rw|sa|sb|sc|sd|se|sg|sh|si|sj|Ja|sk|sl|sm|sn|so|sr|ss|st|su|sv|sx|sy|sz|tc|td|tf|tg|th|tj|tk|tl|tm|tn|to|tp|tr|tt|tv|tw|tz|ua|ug|uk|us|uy|uz|va|vc|ve|vg|vi|vn|vu|wf|ws|ye|yt|yu|za|zm|zw)\b/?(?!@)))'
re=$re_all
if [ $# -gt 0 ] && [ "$1" = '-w' ]; then
re=$re_web
fi
grep -oP "$re" "$@"
- quickthrower2 3y agoThat tld list will be brittle. There is a continuous stream of them coming online.
- networked 3y agoYou are right. I should add a warning. The "Web URL" regular expression has the advantage of catching "example.com/foo?q=bar#baz", not just "https://example.com/foo?q=bar#baz https://example.com/foo?q=bar#baz", but Gruber published it in 2014 (https://gist.github.com/gruber/8891611/revisions https://gist.github.com/gruber/8891611/revisions). I have not updated it. I recommend normally using the script without -w and filtering the output for HTTP(S) URLs. The default regular expression (https://gist.github.com/gruber/249502 https://gist.github.com/gruber/249502) does not rely on a list of TLDs. Edit: Added a warning to the original comment. Thanks for prompting me to. Here is a version without the "Web URL" regex. #! /bin/sh # shellcheck disable=SC1112 set -eu # The URL regular expression by John Gruber. # https://gist.github.com/gruber/249502 re='(?i)\b((?:[a-z][\w-]+:(?:/{1,3}|[a-z0-9%])|www\d{0,3}[.]|[a-z0-9.\-]+[.][a-z]{2,4}/)(?:[^\s()<>]+|\(([^\s()<>]+|(\([^\s()<>]+\)))*\))+(?:\(([^\s()<>]+|(\([^\s()<>]+\)))*\)|[^\s`!()\[\]{};:'"'"'".,<>«»“”‘’]))' grep -oP "$re" "$@"