4 ms·
I feel its the same as with any long standing computer system we have today. It was designed as more and more of the world came online and all the growing pains
by jetbalsa 3y ago
I feel its the same as with any long standing computer system we have today. It was designed as more and more of the world came online and all the growing pains it came with. Could it be built from scratch today better? Yes. Will it? No. I suspect it will be around long after we are all dead. Same with IPv4 :V
- magicalhippo 3y agoFor most software it doesn't really matter either. I've written unicode-aware software for over a decade, doing a wide variety of programs, and I've never had to bother with all that mess. If I'm parsing strings I'm looking for stuff in the 7-bit ASCII range which maps neatly onto the Unicode representations, and so I just need to take care to preserve the rest. The only trouble I've had is that a lot of programmers haven't learned, or don't get, that text encoding is a thing and that it needs to be handled. So they'll hand me an XML they claim is UTF-8 encoded, except that XML header was just copypasta and the actual XML document is encoded in some other system encoding like Windows-1252. Or worse, a mix of both.
- hot_gril 3y agoHonestly I like ipv4 better than v6. I like having a NAT and easy addresses like 192.168.1.3 instead of fe80::210:5aff:feaa:20a2. They didn't need to mess with those things just to expand the address space, like how utf8 didn't require remapping ASCII.
- jrockway 3y agoIPv4.1 should have just had 39 bits, to be written like 999.999.999.999. (I know this wouldn't have actually had much effect, nobody is going to add new routes in the middle of "class A" spaces that already existed, so it would just give those that already had IP addresses more IP addresses. Additionally, people really abuse decimal addresses in horrifying ways; for example, Fios steals 192.168.1.100-192.168.1.150 for its TV service, and that range doesn't really correspond to anything that you can mask off in binary. It only makes sense in decimal, which is not what any underlying machinery uses. They should have given themselves a /26 or something. You get 3 for yourself (modulo the broadcast and gateway address), and they get 1 for TV.)
- deleted 3y ago[deleted]
- hot_gril 3y agoHaving it actually be decimal might've been nice, but at this point people are used to the 1-254 range, and I think the least jarring addition of extra bits would be to simply extend it for the addresses that need them (and not for the ones that don't). So you could have 123.444.3.254 or longer like 123.444.3.254.12.43.
- Dylan16807 3y agoGo ahead and use fe80::3 as a link-local address. For site local, fec0::3. Yeah site-local is discouraged but you can still do it. Or you can slightly misuse fd00::3. You only get those latter 16 hex characters if you explicitly don't want to choose addresses.
- vacuity 3y agoBe the change you want to see in the world. If we're going to make huge breaking changes, might as well do it sooner rather than later.
- jetbalsa 3y agoWith something as large as a end user language format for input, this is a change we ourselves cannot make, just as using another calendar for dates. Just because I want to use the year 2002023 calendar with 29.5 days per month, doesn't make it useful to others or myself really.
- nvm0n2 3y agoI think actually you could. A thought experiment: The problems with Unicode are mostly to do with internal inconsistencies and churn, problems that usually only affect programmers. 1. Different ways to encode the same visually indistinguishable set of characters as code points leading to normal forms, text that compares unequal even when it appears to be identical, the disastrous "grapheme clusters" concept and so on. 2. Many different ways to encode the same sequence of code points as bytes. Not only UTF-32/16/8 but also curiousities like "modified UTF-8". 3. Emoji. A fractal of disasters: 3.a. Updates frequently. Neither Unicode nor software in general was built on the assumption that something as basic as the alphabet changes every year. If you send someone an emoji, can their device draw it? Who knows! In practice this means messaging apps can't rely on the OS system fonts or text handling libraries anymore which is a drastic regression in basic functionality. 3.b. (Ab)uses composition so much it's practically a small programming language, e.g. flags are composed of the two letter country code spelled using special characters. People are represented as as generic person plus skin color patch, families are represented using composed individual people etc. 3.c. Meaning of a character is theoretically specified but can subtly depend on the font used, e.g. people use a fruit emoji in visual puns because of how it looks specifically on Apple devices, so a "sentence" can make no sense if it's rendered with a different font. 3.d. Unbounded in scope. There's no reason the Unicode committee won't just keep adding new pictograms forever. 3.e. Encoded beyond the BMP which in theory every correct program should handle but in practice some don't because nobody except a few academics used characters beyond it much until emoji came along. 3.f. Disagreement over single vs double width chars, can only know this via hard-coded tables, matters for terminals and code editors. Some of these can potentially be cleaned up outside of the Unicode consortium in backwards compatible ways. You could have a programming language that automatically normalized strings to fully composed form when deserializing from bytes, and then automatically folded semantically identical code points together (this would be a small efficiency win for some languages too). You could campaign to build a consensus around a specific normal form, like how UTF-8 gained consensus as a transfer encoding. You could also define a fork of Unicode (using private use areas?) that allocates a single code point to the characters that are unnecessarily using composition today but don't yet have one and then just subset out the concept of composition entirely. Emoji are a big problem. It's tempting to say that these should not be encoded as characters at all. Instead there could be a set of code points that define bounds that contain a tiny binary subset of SVG, enough to recreate the Apple pixel art somewhat closely. Emoji would always be transmitted as inlined vector art. Text rendering libraries would call out to a little renderer for each encoded glyph, using a fast fingerprinting algorithm to deduplicate the bytes to an internal notion of a character. To avoid wire bloat, text can simply be compressed with a pre-agreed zstd or Brotli dictionary that contains whatever images happen to be popular in the wild. At a stroke this would avoid backwards compat problems with new emoji, enabling programs working with text to be upgraded once and then never again, eliminate all the ridiculous political committee bike-shedding over what gets added, let apps go back to using system text support and get rid of the bajillion edge cases that emoji have spewed all over the infrastructure.