37 ms·
Any word on when all functions will be UTF-8 safe?
by poobrains 13y ago
Any word on when all functions will be UTF-8 safe?
- jcampbell1 13y agoOther than "substring", how often do you really run into this problem? I have never had a problem with PHP and UTF-8. Oh, good luck trying to match all unicode punctuation with Python/Javascript. With php it is as simple as preg_match('/\pP/u' I deal with a lot of UTF-8 issues, and while php is super ugly, there are pretty good solutions in cases where it requires ugly hacks in other languages.
- Zancarius 13y agostrlen and strpos are two fairly commonly used functions that come to mind.
- desas 13y agoUse mb_strlen and mb_strpos instead where the string may contain multibyte characters.
- Zancarius 13y agoOr patchwork/utf8 [1] since it handles fallback in the event mb_* isn't installed and can utilize several other libs. [1] https://github.com/nicolas-grekas/Patchwork-UTF8 https://github.com/nicolas-grekas/Patchwork-UTF8
- acabal 13y agoPretty often. Yes you can use mb_* functions everywhere but not all string functions have an mb_* counterpart, and forget just once and you risk blowing up the entire rest of the request. Not to mention times when you have to use a 3rd party library that doesn't bother with mb_*. Native UTF is such an important thing... I know it's tough to implement, but come on guys!
- dhoulb 13y agoI do run into this fairly regularly. The mb_ functions are solid, but I was hitting something the other day where a character, I think it was NBSP, was causing a string to output empty in 5.4 (but was working fine in 5.3). I think there's something that needs to be fixed at a core level, maybe with PHP 6, that just guts how the language deals with multibyte. Even if it means making it backwards incompatible. I'd probably take the opportunity to drop the mb_ functions, namespace the entire language, and make it multibyte by default. Needs doing eventually!
- nikic 13y agoIt's a common misconception. Many people don't understand that the normal string functions are perfectly safe on UTF-8, as long as you don't use hardcoded lengths or offsets. I.e. substr($str, 0, 50) is not safe due to the explicit "50" in there, but substr($str, 0, strpos($str, "foo")) will work correctly on any well-formed UTF-8. If people have encoding issues in PHP it usually just means that they didn't manage to set up their database properly (you know, finding which one of the 10 encoding options in MySQL is the right one ;) From my personal experience I've had a lot more issues with encoding in Python than I had in PHP - exactly because PHP ignores encoding and lets me deal with it.
- jcampbell1 13y agoI agree. Python creates a lot of problems that are non-obvious. Like os.walk('.') works just fine, right up to the point where it silently trashes all unicode file names.
- wvenable 13y agoPHP should keep everything as-is and simply add new data type for unicode strings. It should be utterly and completely incompatible with any current function that accepts strings: $binString = "hello"; $unicodeString = u"Hello"; strlen($binString); // 5 strlen($unicodeString); // Error And then developers can slowly start making functions and methods more unicode aware as necessary. Or make an entirely new string API. Then just have functions to take binary strings and convert them to unicode strings (providing an encoding) and unicode strings to binary strings (also providing an encoding). This would be way safer and simpler than PHP6 or Python 3.
- semerda 13y agoAs I understand in PHP to work with Unicode you need to use Multibyte String Functions since Unicode encodes into 2+ bytes vs the traditional 1 byte. So doing a "substr" as per your example would not work unless you use "mb_substr".
- Joeri 13y agoSorry, but PHP's string handling is abysmal. - strlen can't be used for length checks if you care about length in characters, a problem compounded when you use a real database where field lengths are defined in characters instead of bytes - the only way to iterate char by char instead of byte per byte is to use mb_substr, as there is no trick to make $str[$i] do anything but return bytes. - the string and array API's have inverse ordering of their parameters, which means that after a decade of writing PHP I still can't remember which is which without looking it up - Typing mb_ in front of everything is ugly enough, but it also makes autocompletion tricky, especially since the strings don't have methods. (e.g. $str->pos()) - Speaking of which, since the strings don't have methods, they stick out like a sore thumb in OO code. String and array handling code invariably ends up ugly unless you write everything procedural style (and then you have other issues). - Sort() cannot be made to sort unicode on windows, regardless of which parameters you give it. In fact, the only way to sort unicode on windows is by using the Collator from the intl extension. Part of that is microsoft's fault by not supporting UTF-8 in the windows API's at all, but PHP isn't helping. - If you don't care about windows, the proper way to sort text is first calling setlocale(LC_COLLATE, "en_US.UTF8") and then passing the SORT_LOCALE_STRING argument as second parameter to every call to sort(). Ugly, ugly, ugly. - natsort(), aka "natural sort" cannot be used to sort text like a human would expect, in any context. It always produces invalid results, even for ANSI codepages. (e.g. try to sort resume, rope and résumé) - The use of utf8_decode and utf8_encode is actively harmful in almost all circumstancces. There is never a good reason to use them, since the very rare case where you need them iconv or mb_convert_encoding are better suited. Yet, the PHP documentation doesn't tell you this, causing lots of people to be led astray (as I once was). - Oh yeah, and there are no less than three API's for unicode string handling, the mb_ functions, the iconv_ functions and the grapheme_ functions. What's the difference? I don't know, and I really can't be bothered to read PHP's source to find out. - htmlentities() always requires the parameters ENT_QUOTES, "UTF-8" to do its job securely (well, almost, as it doesn't encode forward slash which OWASP recommends). Unless you use a wrapper, your code is yet again uglified. - The secure way to JSON-encode text is, and I kid you not, json_encode($data, JSON_HEX_TAG|JSON_HEX_APOS|JSON_HEX_QUOT|JSON_HEX_AMP). Try typing that three times in a row, I dare you. - And finally, mysql is by far the worst database for unicode handling, because it cannot sort unicode text according to the standard, at all, no matter what you do. That's not PHP's fault ofcourse, but since I'm bitching... :)