4 ms·
Great post. So the follow-up question is, how much work would be involved in mirroring all that in a native UTF-8 encoded string type? :-) Windows interop, and
by nblumhardt 10y ago
Great post. So the follow-up question is, how much work would be involved in mirroring all that in a native UTF-8 encoded string type? :-)
Windows interop, and 8/16 conversions, would obviously be an expense, but ~half the storage requirements of UTF-16 have to represent a substantial CPU/RAM saving.
Guess it's too much cruft and complexity to introduce into .NET now, but who knows? https://twitter.com/terrajobst/status/717935598904807424 https://twitter.com/terrajobst/status/717935598904807424
- matthewwarren 10y agoYeah I think the legacy interop is the tricky part, see "Why does C# use UTF-16 for strings?" (http://blog.coverity.com/2014/04/09/why-utf-16/#.V1AT9vkguUl http://blog.coverity.com/2014/04/09/why-utf-16/#.V1AT9vkguUl) for example. Interestingly enough Java just implemented compact strings that can be ISO-8859-1/Latin-1 (one byte per character) or as UTF-16 (two bytes per character). See http://openjdk.java.net/jeps/254 http://openjdk.java.net/jeps/254 and https://www.infoq.com/news/2016/02/compact-strings-Java-JDK9 https://www.infoq.com/news/2016/02/compact-strings-Java-JDK9 for more info.