4 ms·
A cheap shot by McIlroy. Knuth was specifically asked to write in the literate style. Word counting is one of those simple, domain-independent problems that l
by bluesnowmonkey 15y ago
A cheap shot by McIlroy. Knuth was specifically asked to write in the literate style.
Word counting is one of those simple, domain-independent problems that lend themselves well to code reuse. It's a rare type of problem. Most tasks presented to a professional software developer could not be solved by a small shell script. A large and unmaintainable one, maybe.
- lsb 15y agoMcIlroy's code lends itself very well to literate coding! Here's a try: # First, tr anslate multiple s queezed occurrences of the c omplement of A-Za-z (non-word-characters) into a line separator tr -cs A-Za-z '\n' | # and then lowercase every word. tr A-Z a-z | # Sort the words with a disk-based mergesort. sort | # Count the unique characters :: [String] -> [(Int,String)] uniq -c | # And do a reverse numerical disk-based merge sort. sort -rn | # And write the first $1 lines and then quit. sed ${1}q Personally, I'd change the last two to be sort -rn -k 1,1 | head -n $1 but that's just bikeshedding. And the parts it's made from are so modular, and so focused, that you can wrap your head around all of the problem, without worrying about how sort works, or how tr expands character ranges. If I gave a "professional software developer" in 2011 a problem to count word frequencies, and I got back 10 pages of Pascal that didn't go much faster than code that fits on a Post-It note for my problem, I wouldn't trust that developer with anything else important.
- gnuvince 15y agoYour code fails for French text.
- lsb 15y agoAs did Knuth's; see below the discussion about that. If you wanted to change it, to split "words" by blanks and punctuation characters, you would have tr -s [:blank:][:punct:] "\n" and then that piece would fit into the pipeline, the rest unscathed. This design is far easier to reason about than a dozen pages of Pascal.
- ralph 15y agoThose globs should be protected from the shell otherwise ./bc is going to alter tr's arguments.
- d0mine 15y agoHow many modern programs comply with Unicode Text Segmentation http://unicode.org/reports/tr29/ http://unicode.org/reports/tr29/ for finding word boundaries or Unicode Collation Algorithm http://unicode.org/reports/tr10/ http://unicode.org/reports/tr10/ for sorting the words.
- sedq 15y agoI think the implication is that many would give you back 10 pages. And that's why I generally don't trust the products of most of today's "professional software developers" for doing anything important. But I trust UNIX and the shell functions I write.
- TheRevoltingX 15y agoHonestly, that looks more like examples in a MAN page rather than explanations of an algorithm.
- ralph 15y agoThe reason for preferring sed 42q instead of head -n 42 is because head didn't used to exist. IIRC it was created at Berkeley and for some years there were many systems that didn't have it. It does seem a bit superfluous when all it could do was head -42 (the -n and other options came later). I still write sed 42q to this day. :)
- epo 15y agoAgreed. Knuth no doubt chose a simple, but non-trivial, program to build from scratch to illustrate the technique. McIlroy responded with an equivalent made up from Lego bricks. Knuth documented what he did. McIroy said "I used this command" n times.
- dwc 15y agoAs Knuth was given his task, McIlroy was given his own which he performed well. I disagree that code reuse only works in rare instances. Rather, code reuse works better in small focused pieces than in large chunks. Utilities such as "sort" get used a lot. The rare part is that a whole solution can be made out of UNIX utilities, not that UNIX utilities are used at all. Being familiar with the utilities available, much of the coding I currently do ends up being domain specific processing called from a shell script, often in a pipeline. Also, sometimes I can answer feature requests from users by showing them how they can actually do it themselves on the UNIX command line. This kind of thing is often forgotten, as it's no longer "development." But code I don't have to write because the user can reuse utilities should certainly count.