3 ms·
perl unicode support is actually pretty good http://dheeb.files.wordpress.com/2011/07/gbu.pdf http://dheeb.files.wordpress.com/2011/07/gbu.pdf
by heeen 13y ago
perl unicode support is actually pretty good http://dheeb.files.wordpress.com/2011/07/gbu.pdf http://dheeb.files.wordpress.com/2011/07/gbu.pdf
- logicallee 13y agoupdated to remove refernece to unicode. the additional context involved with knowing Perl isn't enough IMHO to learn Perl just to parse text. I would do that with a 'more popular' language and its regex library, especially for new users.
- einhverfr 13y agoCareful though. There's a lot of parsing that folks want to use regexps for that don't do that well. Why Perl really shines for text file parsing is that you have a fairly large set of tools as appropriate, which fall into largely three categories: 1. regexps (great for some things, lousy for others. DO NOT use these for parsing HTML) 2. Recursive descent parsers. You could write an HTML parser in one of these if you need to. 3. Dedicated format parsers (CSV, XML, JSON, etc). It's the combination of the three that makes Perl a very good tool for parsing text files. One thing going for Perl is a strong community of people who can help point people to the right tool for the right job. For example, I am in the process of writing a library to do PostgreSQL tuple parsing and serialization. It is sort of like Text::CSV with some differences. In this case, regexps are the right tool. But if I was writing something else, I would probably go with recursive descent.