5 ms·
I'm sometimes believe that full general purpose embracing of unicode for text, with no clean distinction between "machine-friendly" text versus "for human" natu
by ahefner 5y ago
I'm sometimes believe that full general purpose embracing of unicode for text, with no clean distinction between "machine-friendly" text versus "for human" natural language text (in every script since the dawn of time plus every goofy emoji anyone dreams up, with all the complexity these entail) is a major mistake that has lead computing astray. I fear, though, that it is impractical to separate these things, short of entirely shunning the latter, and tempting as it is I can't quite advocate a return to pure ASCII.
- zionic 5y agoI’ll say it, it was a mistake. Every developer who learns swift for the next hundred years will curse the designers the first time they need to do “basic” string processing. It’s gotten slightly better, but it’s still a joke.
- int_19h 5y agoAs a developer whose native language is not English, I'm really glad that Swift has finally forced developers from Western countries (and especially US) contemplate what strings actually are and aren't, and how to properly use them in different contexts.
- bouke 5y agoCare to explain? I believe Swift’s model to be quite sensible for human non-US text, having an API that puts Unicode before ASCII. The bytes are still available if you want to shoot yourself in the foot.
- jiggawatts 5y agoProgramming languages really should have distinct string-like types: identifier -- printable ASCII characters only, an array of 8-bit chars ucs16 -- An array of 16-bit chars for compatibility with Windows, Java, and .NET utf8 -- Normalised, 100% valid UTF-8 with potentially some "reasonableness" constraints text -- Abstraction over arbitrary code pages, including both Unicode and legacy encodings. Languages like Rust kinda-sorta implement this. For example, the PathBuf type internally uses a "WTF8" encoding that is vaguely Utf-8 compatible, but allows the invalid code sequences that can turn up in Win32 system API calls. IMHO that's a good try but not ideal. The back-and-forth conversion is complex, requires temporary buffers, and is slow as molasses for many types of API calls. The ideal would be to have abstractions (traits, interfaces, whatever) that cover all the use-cases. E.g.: it should be possible to test if a 'utf8' string contains an 'identifier' string. It should be possible to compare strings without having to convert their formats. Etc...
- mjevans 5y agoMy take on the 'string like types' languages should have: # Simple indexed access (often an array, possibly an array of arrays) raw / octets -- Not-classified sequence of raw bytes # Fancy strings, which MAY be validated (but don't have to be), and MAY stay validated (if the operations are known to be simple enough), and MAY also have multiple types of index for speedy access to specific points by raw byte, unit run of encoded components, complete display units (a single displayed element), or even a cached last known render left bound on an output. UTF-8 / WTF8 / UTF8 / ASCII -- A fancy string with octet components of possibly multi-byte (variable) length UCS-2 / USC-16 / etc -- A legacy string encoding format that no one should use as it too is variable length but suffers from endien confusion in raw data storage. In an object based language the latter two would probably use the first as a raw storage mechanism for the strings, while they'd also have some associated attributes for the desired encoding, if it's validated, how it is known to be normalized, and storage for different index aids. Crucially the system libraries should have the same interface. If there isn't library support for converting / normalizing encoded text the results should always be raw octets. If library support is included then WTF8 should probably be the result target of any operations, possibly upgraded to UTF-8 if the results are theoretically still valid.
- mlindner 5y ago> ucs16 -- An array of 16-bit chars for compatibility with Windows, Java, and .NET Just no. Maintaining compatibility with broken implementations that hold incorrect assumptions is not something that should be done. Especially not encoding it into some kind of standard.
- jiggawatts 5y ago"Should not be done" will often lose out to the more pragmatic backwards-compatible solution. Those three platforms describe 90-95% of all "enterprise" business software ever written or the platform they're used on. Disregarding that weight of history for... what? A clever trick that Linux used to finally add i18n support decades after other platforms?