6 ms·
Stringref is an extremely thoughtful proposal for strings in WebAssembly. It’s surprising, in a way, how thoughtful one need be about strings. Here is an aside
by paroneayea 3y ago
Stringref is an extremely thoughtful proposal for strings in WebAssembly. It’s surprising, in a way, how thoughtful one need be about strings.
Here is an aside, I promise it’ll be relevant. I once visited Gerry Sussman in his office, he was very busy preparing for a class and I was surprised to see that he was preparing his slides on oldschool overhead projector transparencies. “It’s because I hate computers” he said, and complained about how he could design a computer from top to bottom and all its operating system components but found any program that wasn’t emacs or a terminal frustrating and difficult and unintuitive to use (picking up and dropping his mouse to dramatic effect).
And he said another thing, with a sigh, which has stuck with me: “Strings aren’t strings anymore.”
If you lived through the Python 2 to Python 3 transition, and especially if you lived through the world of using Python 2 where most of the applications you worked with were (with an anglophone-centric bias) probably just using ascii to suddenly having unicode errors all the time as you built internationally-viable applications, you’ll also recognize the motivation to redesign strings as a very thoughtful and separate thing from “bytestrings”, as Python 3 did. Python 2 to Python 3 may have been a painful transition, but dealing with text in Python 3 is mountains better than beforehand.
The WebAssembly world has not, as a whole, learned this lesson yet. This will probably start to change soon as more and more higher level languages start to enter the world thanks to WASM GC landing, but for right now the thinking about strings for most of the world is very C-brained, very Python 2. Stringref recognizes that if WASM is going to be the universal VM it hopes to be, strings are one of the things that need to be designed very thoughtfully, both for the future we want and for the present we have to live in (ugh, all that UTF-16 surrogate pair pain!). Perhaps it is too early or too beautiful for this world. I hope it gets a good chance.
- tomcam 3y agoI think the WebAssembly people have been judicious about features. Watching it evolve has made me feel that they truly respect how important it is to keep things well thought out and as efficient as possible. I feel like it’s in very good hands.
- paroneayea 3y agoI agree with this assertion. WebAssembly, on its whole, is extremely good. The string stuff is, IMO, something the group has not come to realize the "right direction" on, but so much has been done right! Hopefully strings can get there too. :)
- pjmlp 3y agoIt guess that is why GC support in now in about 5 years and counting, whereas CLR is doing it since 2001, including with interoperability with C++.
- nicoburns 3y ago> right now the thinking about strings for most of the world is very C-brained, very Python 2 Is it? Doesn't pretty much every language have a unicode string type (be that UTF16 in older languages or UTF8 in newer one) that is the default goto type for dealing with text these days? C and C++ being the notable exceptions I suppose.
- kragen 3y agoutf-8 strings work fine in c and c++, as they have since utf-8 was introduced; that was the major design objective of utf-8 in fact
- nine_k 3y agoThey work well as long as you're fine working with bytes. For "characters" which a user sees on the screen, that is, graphemes, you need an entirely new layer. Take some word, e.g. "éclair". How long is it? What are its first three characters? How do you uppercase it?
- kragen 3y agothose are library functions, and they work fine on utf-8 strings, though graphemes in particular are difficult and context-dependent in unicode in a way that is exactly the same in c and in java
- amluto 3y agoThey are library functions for which a good library does not exist. I recently needed to convert probably-UTF-8 data to definitely valid UTF-8 with errors replaced. This was not an enjoyable experience in C++. (The ztd proposal is IMO a big step in the right direction.)
- flohofwoe 3y agoStuff like this is handled in a (3rd-party) UNICODE library in the C/C++ world, which should ideally work on UTF-8 encoded byte arrays, provided by another (3rd-party) UTF-8 encoding/decoding library. Other then that high-level UNICODE stuff (like finding grapheme cluster boundaries) UTF-8 itself really works fine in C/C++ anywhere than Windows in the sense that I can write a foreign-language "Hello World!" and it "just works" (e.g. if the whole source file is UTF-8 encoded anyway, than C string literals are also automatically valid UTF-8 strings). UNICODE on Windows is still a bigger mess than it should be because of its UCS-2 / UTF-16 heritage.
- kragen 3y ago> Python 2 to Python 3 may have been a painful transition, but dealing with text in Python 3 is mountains better than beforehand it is not python 2 made a disastrously wrong choice about how to add unicode support python 3 inserted that disastrously wrong choice everywhere (though at least you no longer get compile errors when you put a non-ascii character in utf-8 or latin-1 in a comment, a level of brain damage i've never seen from any other language) rust and golang made reasonable choices about how to handle unicode; python, by contrast, is a bug-prone mess i've lost python error tracebacks generated by an on-orbit satellite because they contained a non-ascii character and so the attempt to encode them as text generated an encoding error. python's unicode handling catastrophe has made it unusable for any context where reliability is especially important
- amluto 3y agoI would argue that Python 3 reliability issues should be blamed on inadequate static checking, not on Unicode strictness. If you do foo.decode(), you are introducing an operation that can throw. If you are programming in Python for a reliability-critical environment, you should detect this at commit/test time and handle it appropriately. Rust is every bit as Unicode-strict, but it’s harder to fail to notice that you have a failure path. Meanwhile, Python 2 will just happily malfunction and carry on. Sure, the code keeps executing, but this doesn’t mean that you will actually get your error message out.
- kragen 3y agopython has a ubiquitous lack of static checking; every other feature added to it must be considered in that context. if on balance it's bad without static checking, it's bad in python the code in question was not doing foo.decode() or foo.encode(). it was writing a string to a file. python 3 inserts implicit unicode encoding and decoding operations in every access to environment variables, file names, command line arguments, and file contents, unless you pass a special binary flag when you open the file, as if you were on fucking ms-dos. all those things are byte strings, and rust and python 2 give you access to them as byte strings. python 3 instead prefers to insert subtle bugs into your program
- 3y ago
- flohofwoe 3y agoPython3 is really not a great example to copy elsewhere though. By the time Python3 came about it was already clear that UTF-8 encoding is all one ever needs to represent UNICODE strings, and all the other encodings are either historical accidents (like UCS-2 and UTF-16), or only needed at runtime in very specific situations (like UTF-32, but even this is debatable when working with grapheme clusters instead of codepoints). And with that basic idea that strings are just a different view on a bytestream (e.g. every string is a valid bytestream, but not every bytestream is a valid string) most of the painful python2-to-python3 transition could have been avoided. I really don't know what they've been thinking when the 'obviously right' solution ("UTF-8 everywhere") was right there in plain sight since around the mid-90's.
- amluto 3y ago> And with that basic idea that strings are just a different view on a bytestream (e.g. every string is a valid bytestream, but not every bytestream is a valid string) most of the painful python2-to-python3 transition could have been avoided. Can you elaborate? Much of the pain of the transition was figuring out which strings were bytes and which were Unicode data. The actual spelling of the type names never seemed like a big deal to me. (I do think Python 3 messed some things up. My current favorite peeve is the fact that iterating bytes yields ints. That causes a lot of type confusions to result in digit gobbledygook instead of a useful exception or static checker error.)
- altfredd 3y agoUnless you enjoy getting hacked, all strings received from outside sources are bytes.
- flohofwoe 3y ago> Much of the pain of the transition was figuring out which strings were bytes and which were Unicode data. And for a lot of code (that which just passes data around), this shouldn't matter. It's basically "Schroedinger's strings", you don't need to know if some data is valid string data until you actually need it as a string, and often this isn't needed at all (IMHO all encodings/decodings should be explicit, not just between bytestreams and strings, but also between different string encodings - and those should arguably go into different string types which cannot be assigned directly to each other - e.g. the standard string type should always only be UTF-8). Also, file operations should always work on bytestreams (same in the IO functions of the C stdlib btw).
- flir 3y agoVery early in my career, I said something about strings and a more experienced programmer said "that's because you think a string is an array of bytes terminated with a \0". Absolute lightbulb moment for me, and not just about strings.