4 ms·
Except these are all well outside the ambit of what programmers usually think of as text processing, so they won't try to solve them using the same tools. More
by derleth 13y ago
Except these are all well outside the ambit of what programmers usually think of as text processing, so they won't try to solve them using the same tools.
More to the point, they sound hard, so people won't be so quick to claim they've solved them.
On the other hand, case-insensitive string matching sounds easy, even if it's actually somewhat difficult due to the language dependencies mentioned above, so people will claim to have a general solution that fails the first time it's faced with i up-casing to İ instead of I, or the fact the German 'ß' up-cases to 'SS' as opposed to any single character. (Unicode does contain 'ẞ', a single-character capital 'ß', which occurs in the real world but is vanishingly rare. As far as modern German speakers are concerned, the capital form of 'ß' is 'SS'.)
http://en.wikipedia.org/wiki/Capital_%E1%BA%9E http://en.wikipedia.org/wiki/Capital_%E1%BA%9E
http://opentype.info/blog/2013/11/18/capital-sharp-s-design-review/ http://opentype.info/blog/2013/11/18/capital-sharp-s-design-...
http://blogs.msdn.com/b/michkap/archive/2009/07/28/9850675.aspx http://blogs.msdn.com/b/michkap/archive/2009/07/28/9850675.a...
http://www.personal.psu.edu/ejp10/blogs/gotunicode/2008/07/a-new-german-unicode-letter-ca.html http://www.personal.psu.edu/ejp10/blogs/gotunicode/2008/07/a...
- taeric 13y agoRight, I do not disagree. I just feel better treating them the same. That is, both are actually easy and reliable so long as you realize you have to make some gross simplifications. And most of the time your life will be much easier if you start with the gross simplifications and try to expand beyond them only when necessary. (This is also why I'm loathe to try programming in unicode...)