4 ms·
> Perl6 is still over 9 times slower than Perl5 Any chance you can share what code, data, and compiler version you used to get those numbers? > Until Perl6's
by raiph 7y ago
> Perl6 is still over 9 times slower than Perl5
Any chance you can share what code, data, and compiler version you used to get those numbers?
> Until Perl6's performance at least matches Ruby's for basic parsing (no, not grammars) it's not going to be taken seriously no matter how many other tricks it has up its sleeve.
I agree it's still slow but question your prognosis. For example, I think devs who care about O(1) handling of graphemes will use it when it's fast enough for them. And its multi prompt continuation control flow/concurrency/parallel/async design is bearing fruit.
But even if you're right, I see no reason not to think that P6 will keep getting faster and develop nicely throughout the 2020s.
- cutler 7y agoPerl 5.26.3, Ruby 2.6.3, Rakudo Star 2019.03.1 all running on OS X 10.12.6 use v6; for 'logs1.txt'.IO.lines -> $line { say $line if $line ~~ m/<<\w ** 15>>/; }
- raiph 7y agoThanks! One fly in the ointment is that you didn't share anything about the data, especially its encoding and what sort of multilingual text it contained, nor the code for the other proglangs. One upshot is that I'm now wondering if you're comparing apples with oranges. P6 is optimized for Unicode, making handling of Unicode simple, with O(1) character (grapheme) indexing. For example, in P6 regex syntax, << and \w, which are related to the concept "word character", refer to graphemes ("what a user thinks of as a character") whose base codepoint has the "word" property. And \w 15 refers to 15 of these graphemes. In contrast \w in P5 regex syntax, by default, refers to a codepoint with the "word" property, not a grapheme. And I'm not sure what support Ruby has for graphemes. So I'm wondering if your log text was utf8, and if so, if you took the codepoint/grapheme distinction into account.
- cutler 7y agoThe data was a 19Mb web server log file so pretty standard stuff. I've heard these sort of excuses for Perl 6's sluggishness before. Whatever exalted abilities Perl 6 may have it's pretty meaningless if that comes at the expense of performance on more basic tasks. Besides, modern Rubies use utf-8 by default so the comparison still stands. I'd also be surprised if the latest Perl 5 wasn't also unicode compatible.
- raiph 7y ago> The data was a 19Mb web server log file so pretty standard stuff. Then it should be treated as likely being Unicode text, likely utf8, and possibly containing text of any language supported by Unicode, unless you expressly know otherwise. In which case, unless you specifically ensured your regexes were written to process characters as graphemes then they're incorrect. Of course, if your logs only contain English text or text in other languages whose alphabets only use one codepoint per character then they'll have worked. But so what? > I've heard these sort of excuses for Perl 6's sluggishness before. Have you understood them? Are you saying you defend fast, if only English people visited a website, but incorrect if Indian people did, over sluggish but correct no matter who visits a website? > Whatever exalted abilities Perl 6 may have it's pretty meaningless if that comes at the expense of performance on more basic tasks. Caring about performantly doing basic things without caring how badly they're done is wrong headed. What I'm speaking of here isn't some exalted thing. It's about absolute basics and being correct: C=G, Character=Grapheme, per the Unicode standard. Just as an ASCII byte is no longer the valid character unit of a web server log, so too a codepoint isn't either. Apple, and hence Swift, gets this right, because they understood the need to handle text correctly. In the meantime, most others are using proglangs (and regex engines) designed before there was widespread awareness of how Unicode actually works. (It's incredible but true that the latest Python 3 doc still doesn't even mention the word "grapheme". This contrasts incredibly sharply with P5 doc which has included discussion of the impending trainwreck of graphemes for over 2 decades but unfortunately can't dig out from the sheer complexity of dealing with it in a backwards compatible way. Hence the infamous https://stackoverflow.com/a/6163129/1077672 https://stackoverflow.com/a/6163129/1077672 Still, I'd rather have P5's honesty and correctness than Python 3's vapid head-in-sand burying.) > Besides, modern Rubies use utf-8 by default so the comparison still stands. I'd also be surprised if the latest Perl 5 wasn't also unicode compatible. The issue isn't about being Unicode compatible for an entire string. (Which for your log files could be one line.) It's about characters and substrings (and regex pattern matches). There is no issue if your web server log only contains English. Likewise if it only contains text in other simple languages for which C=C, Character=Codepoint. But how do you know your log doesn't contain, say, Indian text? Even though P5 and Ruby are Unicode compatible at the level of entire strings, and P5 has one of the best legacy regex implementations for dealing with graphemes, it's still the case that if you write regexes without explicitly taking graphemes into account, then those regexes will fail when applied to Unicode text that contains C=G text such as Indian text.