3 ms·
There is another program I use for editing that is older than ed. It is written in asm. I think it may actually be faster than sed (and sed is faster than AWK,
by textmode 8y ago
There is another program I use for editing that is older than ed.
It is written in asm.
I think it may actually be faster than sed (and sed is faster than AWK, Lua, Perl, Python, etc.)
1.spt:
; x = " - baz"
; y = " - elephant"
;a a = input :f(end)
; output = a
; a ? x :s(d)f(a)
;d output = y
; :(a)
;end
spitbol 1.spt < foo
- davidgould 8y ago>sed is faster than AWK Depends on the awk implementation and the task. However even gnu awk (gawk) is very fast and mawk is astonishing. Here is a simple example: count the lines, words, and characters in a 65MB text file (10 copies of a novel stuck together). Testing on Ubuntu GNU/linux 16.10 reporting middle of three tries: export LANG=ASCII # avoid differences due to unicode $ time -p wc big10.txt 1284570 10956950 64886660 big10.txt real 0.29 user 0.28 sys 0.01 $ time -p gawk '{l+=1; w+=NF; c+=length($0)+1} END {print l, w, c}' big10.txt 1284570 10956950 64886660 real 0.55 user 0.53 sys 0.01 Not bad, gawk is less than twice as slow as wc which is the standard tool for this. $ time -p mawk '{l+=1; w+=NF; c+=length($0)+1} END {print l, w, c}' big10.txt 1284570 10956950 64886660 real 0.35 user 0.33 sys 0.01 But mawk is only 20% slower than wc. For a script! Just for a check, even python is not terrible at this: #!/usr/bin/python import sys l, w, c = 0, 0, 0 for line in file(sys.argv[1], "rb"): l += 1 w += len(line.split()) c += len(line) print l, w, c $ time -p ./wc.py big10.txt 1284570 10956950 64886660 real 0.87 user 0.86 sys 0.01 About 3 times slower than wc and mawk.
- textmode 8y agoHow can I download big10.txt or the novel to recreate it?
- davidgould 8y agobig10.txt is just 10 copies of big.txt from the Peter Norvig spelling corrector essay [0]. [0] http://www.norvig.com/big.txt http://www.norvig.com/big.txt
- textmode 8y agoOn a much slower computer... time -p wc big10.txt 1284570 10956950 64886660 big10.txt real 2.76 user 2.68 sys 0.08 Trying this as novice with k3. Because novice, 2 out of 3 counts are incorrect and probably not the fastest solution used. Total "words" in the example was simply AWK's NF. But looking at big10.txt there anomalies such as words separated by "--" instead of space. Here I used non-space character followed by space. Far from accurate but not too far. 1.k: w:0:"big10.txt";v:,/$w m:v _ss "[^ ] " / "word": char followed by space #w / lines 1+#m / words #v / characters time -p k 1 1284570 10019630 63602090 real 2.70 user 2.40 sys 0.28 Counting lines with sed time -p wc -l big10.txt 1284570 big10.txt real 0.13 user 0.06 sys 0.07 sed -n '$!d;=' big10.txt 1284570 real 0.29 user 0.19 sys 0.09
- davidgould 8y agoThat is a slow computer, mine is a pre-haswell i3. $ time -p sed -n '$!d;=' big10.txt 1284570 real 0.07 user 0.06 sys 0.00 time -p mawk 'END {print NR}' big10.txt 1284570 real 0.04 user 0.03 sys 0.00 $ time -p gawk 'END {print NR}' big10.txt 1284570 real 0.14 user 0.13 sys 0.00 $ time -p wc -l big10.txt 1284570 big10.txt real 0.02 user 0.02 sys 0.00
- textmode 8y agoRevised 1.k. w:0:"big10.txt";v:{" ",x}'w;u:{#v[x] _ss " [^ ]"}'!#v;t:{#w[x]}'!#w #w / lines +/u / words +/t / chars Counts for words and chars are closer but still short due to inexperience using k. But it appears the script is now faster than wc. time -p wc big10.txt 1284570 10956950 64886660 big10.txt real 2.78 user 2.66 sys 0.12 time -p k 1 1284570 10956830 63602090 real 2.57 user 2.42 sys 0.14
- textmode 8y agoHere is how spitbol script measures against wc. As with k, I am lacking in spitbol experience and so the counts are not identical to wc. Also I am using 10MB of big10.txt instead of the entire file. 1.spt: ;* m line count, c word count, o char count ;* p word pattern ; n = "0123456789" ; w = n &ucase &lcase "-" ; p = break(w) span(w) ;a a = input :f(c) ; o = o + size(a) ; m = m + 1 ;b a ? p = :f(a) ; c = c + 1 :(b) ;c output = m ' ' c ' ' o ;end dd if=big10.txt bs=5m count=2 of=10m.txt time -p wc 10m.txt 201346 1763181 10485760 10m.txt real 0.51 user 0.44 sys 0.00 time -p spitbol 1.spt < 10m.txt 201347 1770831 10235302 real 0.33 user 0.31 sys 0.01 It appears that spitbol script is faster than wc.