10 ms·
Joe Armstrong: "In my opinion Erlang is brilliant at handling text"
- bugs 17y agoI don't know very much about erlang but I can understand why it has the stigma that it doesn't handle text well as almost every introduction I have seen has said this is the case due to it being created by telecommunication companie[s]. However my familiarity is limited and whether this is true or not I cannot comment on, but comparing erlang to C in performing text based operations is probably silly if the other person has an option such as say perl available to them.
- mahmud 17y agoString processing is just one of those things you can't offload to a second process. Having Perl in the same box doesn't mean you have the luxury to open some IPC channel and pass your texts to Perl. String processing is one of those types of problems that just benefit from a little thought and 10 seconds of planning. C itself is bogged down by its own horrible string representation, the nul termination, where most operations need to traverse the string up to the terminator before they know where it ends. This causes all sorts of horrible buffering tricks, conditional tests on every character, and other unpleasant things. Pascal had length-prefixed strings from day one, and DJB created Netstrings to make network programming easier on people. See http://c2.com/cgi/wiki?LeasedString http://c2.com/cgi/wiki?LeasedString
- cloudhead 17y agoEven though he has a good point: strings are lists, and erlang is good at lists — erlang doesn't have native regexps like perl, js, or ruby, nor is the support for them that good.
- ellyagg 17y agoAside from the fact that regex are implemented as a library, can you elaborate how support for them is not that good? I think that's wrong: http://www.erlang.org/doc/man/re.html http://www.erlang.org/doc/man/re.html
- rjurney 17y agoThis is true - looked at making a very concurrent webcrawler in Erlang, and the regex bit was painful.
- babo 17y agoFrom 12Bx that changed, regexps are enjoyable but still not as rich as Perl.
- kscaldef 17y agoHmm... I also wrote a concurrent webcrawler in Erlang, and at no point was I tempted to use regular expressions.
- rjurney 17y agoThen we were crawling for different purposes - imagine that! :)
- aaronblohowiak 17y agoinstead of being snarky, can you please provide the reason why regular expressions were required?
- rjurney 17y agoI needed to do pattern matching to pull data out of many different formats, then clean the data before processing it. We were parsing radio station song feeds, and the data is varied, chaotic and often unavailable. This was also a port of a perl POE app, and so it was regex intensive in its original implementation and I don't see how you could effectively deal with such noisy text without regexes. As to snark - his reply was snark, I just replied in kind. Regexes are incredibly useful for all kinds of things, most especially in parsing data from web services. The fact that he didn't need them isn't 'funny,' it means we were doing different things. In any case - using Erlang for something like this is so much win. The POE, and threaded implementations got real ugly real fast as we scaled it up. I knew Erlang could do it - across boxen, without a problem. It sounds like the regex libs have improved, and I look forward to using Erlang again in the future. Happy? :D
- mahmud 17y agoLisp doesn't have native regexps, but of the two main Lisp regexp libraries, one, cl-irregsexp, is 3 times faster than Perl, 6 times faster than Python, 7 times than C, and 8 times faster than Ruby. Scroll to the bottom of the page http://common-lisp.net/project/cl-irregsexp/ http://common-lisp.net/project/cl-irregsexp/ If you have been following some recent papers on regexp performance, there is consensus that things could be a lot faster with better algorithms. I expect the game to change dramatically soon.
- alxv 17y agoThe cited benchmark is completely bogus. The benchmark compares the different implementations based a single trivial regular expression: /indecipherable|undecipherable/. You simply cannot claim a regex engine is faster than another with such a poor experiment. It is evident that Boyer–Moore string search algorithm will outshine any engine on that regular expression.
- mahmud 17y agoOuch! Ineed, it's a lousy benchmark; I only recommended it from memory because it blew me away the first time I saw it. FWIW, if anybody can recommend a good benchmark, I would be happy to do a write up since I am proud of the regex performance of the other CL library (cl-ppcre.)
- jerf 17y agoFrankly, I think "native regexps" has proved a mistake, not a virtue. If you're using so many regexps in your code that you actually care whether the syntax is optimized for it, the odds of you Doing It Wrong (TM) are very, very high. (And it has to be direct use, too... if you want something like a regexp-based dispatch map like Django uses, Perl doesn't even have a significant character advantage over Python; r"" vs qr//.) Making it easy to do the wrong thing and harder to do the right thing (than the wrong thing) has certainly wrecked up a lot of Perl code I've had to deal with. The majority of my professional programming is in Perl with many other programmers, and the number of times I see people do something like $settings =~ /read/ over a string containing settings, often complete with more than one setting that has the substring "read" even though they're only looking for one particular one... oi. Makes me sick.
- aaronblohowiak 17y agoi think this is a cultural thing. ruby is at least as convenient to deal with regexes as perl, but we don't have the same kinds of problems. i think this is because in ruby, the culture is to use the hippest tool for the job. frankly, regular expressions aren't very hip. YAML or JSON is much more hip and therefore likely to fill that niche. Perl-users made regular expressions something that we all expect each other to know. Did they go too far? Seems that way (just like java-users obsession with introducing abstraction layers helped the programming community learn DP.) Your point implies that exposing a powerful mechanism that is easily abused should be avoided. I disagree. I think this is a matter of culture and not "law".
- adamc 17y agoPicking what approach to use based on whether it is "hip" is just a terrible way to write software.
- dtf 17y agoHow does Erlang fare with Unicode in practice?
- zacharypinter 17y agoIf I understand it correctly, Erlang fares well here because it stores each character as a 32-bit integer in a list. The implication of that approach is more memory overhead for each character, but that allows you to treat strings as a list of characters and gives you all the benefits of the standard list-manipulation libraries.
- naz 17y ago> Erlang fares well here because it stores each character as a 32-bit integer in a list Which is silly because if you make a list of 32 bit integers that have ASCII equivalents then Erlang assumes it is a string.
- tlack 17y agoYeah, but still: it lets you operate on unicode "characters" pretty easily. With a ton of associated downsides (unclear string/list of integers dichotomy, wasted memory for ascii, etc)
- silentbicycle 17y agoThe main difference is the way it's displayed by default in the shell. Strings are handled with list operations, and Erlang is great for list operations.
- anonjon 17y agoErlang doesn't make any assumptions about what it is. To Erlang, it is a list of 32 bit integers. There is no difference between a list of integers and a list of characters. The shell will transform this list of 32 bit integers with ASCII equivalents into text for your convenience, but it doesn't do any sort of conversion or typing.
- alxv 17y ago
- ellyagg 17y agoErlang now has a full-featured standard regex module called re that handles utf8 strings and binaries. It now has a standard module for handling unicode. Yes, regexes aren't part of the language syntax, but then that's not true for, say, python, either. Having to use a library to do regex is not going to be the difference maker in productivity for your app. Besides, while I love regex, I still use them as a last resort. You're going to want to use a real parser for reliable structured text processing. Generally, erlang programmers keep strings in binaries, which are compact. Most modules for handling string type tasks allow you to do this, e.g., the re module understands a unicode_binary type. In 6 months of programming erlang professionally, in a domain dominated by scripting languages like python and ruby, I've certainly never been tempted to bolt over string handling issues. Erlang's flexible distribution, concurrency, and reliability model is just too compelling. To get competitive performance in a reasonable amount of programmer time for concurrent applications in other languages you're limited to the subset of tasks that, say, twisted or tornado makes easy. The program I'm writing now couldn't have been done with either of them. Frankly, if it's the choice between built-in support for regex and built-in primitives for distribution, concurrency, and fault tolerance, there's no question in my mind which is more important.
- amix 17y agoHe compares Erlang's string handling with C... And C sucks at string handling. He should compare Erlang to all the modern languages where strings are a first class citizen. Comparing Erlang's string handling to Java, Python or Ruby might change his mind.
- deleted 17y ago[deleted]
- shiro 17y agoI think the main (and almost the only) reason to have distinct string type is performance, not for the ease of programming. Many string operations are useful as a general list operations as well, including regular expression matcher. (Note: Some people emphasize importance of O(1) access of string access by index, but using integer index is also a performance hack. If search operations can return some way to point to the substring you don't need integer indexes.) Another minor reason is to display; people prefer reading sequence of characters in a string syntax. If you have a statically typed language it is easy to display a list of chars in string syntax instead of list syntax. For a dynamically typed language with heterogeneous lists, it can be a performance penalty to check whether a list entirely consists of characters or not at runtime. So, in a sense, it is also about a performance. (Note: Having a syntax for strings has nothing to do with having distinct type for strings. The string syntax can be just a syntax sugar.) But performance is important, of course. One thing very common in string (a list of characters) but not very common in general lists is concatenation. To be precise, lazy language programmers use list concatenation without a guilt, but eager language programmers tend to avoid it since it may cause unnecessary copying of lists. So for the eager evaluation languages, it makes sense to have a string type that has very cheap concatenation operation (e.g. using tree representation) internally.
- amix 17y agoPerformance and ease of use are the main reasons why strings should be first class citizens. Strings are one of the most used data structures and most of today's popular languages have very good support for them, encodings of them and manipulation of them. C does not have that good support for them. Ignoring strings and labeling them as "a list of integers" is a step backward, since a lot of the data we have today is textual and will continue to be textual in the future.