24 ms·
We'd be better off with 9-bit bytes
- FrankWilhoit 1y agoThat's what the PDP-10 community was saying decades ago.
- Keyframe 1y agoYeah, but hear me out - 10-bit bytes!
- Waterluvian 1y agoUh oh. Looks like humanity has been bitten by the bit byte bug.
- pdpi 1y agoOne of the nice features of 8 bit bytes is being able to break them into two hex nibbles. 9 bits breaks that, though you could do three octal digits instead I suppose. 10 bit bytes would give us 5-bit nibbles. That would be 0-9a-v digits, which seems a bit extreme.
- pratyahava 1y agoCrockford base32 would be great. it is 0–9, A–Z minus I, L, O, U.
- pdpi 1y agoThe moment you feel the need to skip letters due to propensity for errors should also be the moment you realise you're doing something wrong, though. It's kind of fine if you want a case insensitive encoding scheme, but it's kind of nasty for human-first purposes (e.g. in source code).
- xpe 1y ago> The moment you feel the need to skip letters due to propensity for errors should also be the moment you realise you're doing something wrong, though. When you think end-to-end for a whole system and do a cost-benefit analysis and find that skipping some letters helps, why wouldn't you do it? But I'm guessing you have thought of this? Are you making a different argument? Does it survive contact with system-level thinking under a utilitarian calculus? Designing good codes for people isn't just about reducing transcription errors in the abstract. It can have real-world impacts to businesses and lives. Safety engineering is often considered boring until it is your tax money on the line or it hits close to home (e.g. the best friend of your sibling dies in a transportation-related accident.) For example, pointing and calling [1] is a simple habit that increases safety with only a small (even insignificant) time loss. [1] https://en.wikipedia.org/wiki/Pointing_and_calling https://en.wikipedia.org/wiki/Pointing_and_calling
- pdpi 1y agoYou misunderstood me. I started off by saying that 0-9a-v digits was "a bit extreme", which was a pretty blatant euphemism — I think that's a terrible idea. Visually ambiguous symbols are a well-known problem, and choosing your alphabet carefully to avoid ambiguity is a tried and true way to make that sort of thing less terrible. My point was, rather, that the moment you suggest changing the alphabet you're using to avoid ambiguity should also be the moment you wonder whether using such a large number base is a good idea to begin with. In the context of the original discussion around using larger bytes, the fact that we're even having a discussion about skipping ambiguous symbols is an argument against using 10-bit bytes. The ergonomics or actually writing the damned things is just plain poor. Forget skipping o, O, l and I, 5 bit nibbles are just a bad idea no matter what symbols you use, and this is a good enough reason to prefer either 9-bit bytes (three octal digits) or 12-bit bytes (four octal or three hex digits).
- int_19h 1y agoClearly it should be 12 bits, that way you could use either 3 hex digits or 4 octal ones. ~
- monocasa 1y agoAlternate world where the pdp-8 evolved into our modern processors.
- tzs 1y ago10-bit has sort of been used. The General Instrument CP1600 family of microprocessors used 16-bit words but all of the instruction opcodes only used 10 bits, with the remaining 6 bits reserved for future use. GI made 10-bit ROMs so that you wouldn't waste 37.5% of your ROM space storing those 6 reserved bits for every opcode. Storing your instructions in 10-bit ROM instead of 16-bit ROM meant that if you needed to store 16-bit data in your ROM you would have to store it in two parts. They had a special instruction that would handle that. The Mattel Intellivision used a CP1610 and used the 10-bit ROM. The term Intellivision programmers used for a 10-bit quantity was "decle". Half a decle was a "nickel".
- ajuc 1y ago5 bit nibbles could just be Baudot Code (A-Z + some control characters).
- ochrist 1y agoCame here to say that. But to my knowledge there has never been any computers based on 5 bits.
- pratyahava 1y agodeleted
- relevant_stats 1y agoI really don't get why some people like to pollute conversations with LLMs answers. Particularly when they are as dumb as your example. What's the point?
- svachalek 1y agoSame, we all have access to the LLM too, but I go to forums for human thoughts.
- pratyahava 1y agook, agree with your point, i should have got the numbers from chatgpt and just put them in the comment with my words, i was just lazy to calculate how much profit we would have with 10-bit bytes.
- pratyahava 1y agoumm, i guess most of the article is made by llm, so i did not see it as a sin, but for other cases i agree, copy-pasting from llm is crap
- CyberDildonics 1y agoOn hacker news the comments need to be substance written by a person, but the articles can be one word title clickbait written by LLMs.
- iosjunkie 1y agoNo! No, no, not 10! He said 9. Nobody's comin' up with 10. Who processing with 10 bits? What’s the extra bit for? You’re just wastin’ electrons.
- ajuc 1y agoHaving 0-1000 fit inside a byte would be great. 1 byte for tonnes, 1 for kg, 1 for g, 1 for mg. Same with meters, liters, etc. Would work great with metric. Or addressing 1 TB of memory with 4 bytes, and each byte is the next unit: 1st byte is GB, 2nd byte is MB, 3rd byte is KB, 4th byte is just bytes.
- titzer 1y ago10 bit bytes would be awesome! Think of 20 bit microcontrollers and 40 bit workstations. 40 bits makes 5 byte words, that'd be rad. Also, CPUs could support "legacy" 32 bit integers and use a full 8 bits for tags, which are useful for implementing dynamic languages.
- phpnode 1y agowhy stop there? 16-bit bytes would be so much cleaner
- dboreham 1y agoUTF-16 enters the chat..
- zamadatix 1y agoBecause we have 8 bit bytes we are familiar with the famous or obvious cases multiples-of-8-bits ran out, and those cases sound a lot better with 12.5% extra bits. What's harder to see in this kind of thought experiment is what the famously obvious cases multiples-of-9-bits ran out would have been. The article starts to think about some of these towards the end, but it's hard as it's not immediately obvious how many others there might be (or, alternatively, why it'd be significantly different total number of issues than 8 bit bytes had). ChatGPT particularly isn't going to have a ton of training data about the problems with 9 bit multiples running out to hand feed you. It also works in the reverse direction too. E.g. knowing networking headers don't even care about byte alignment for sub fields (e.g. a VID is 10 bits because it's packed with a few other fields in 2 bytes) I wouldn't be surprised if IPv4 would have ended up being 3 byte addresses = 27 bits, instead of 4*9=36, since they were more worried with small packet overheads than matching specific word sizes in certain CPUs.
- marcosdumay 1y agoWell, there should be half as many cases of multiples-of-9-bits ran out than for multiples-of-8-bits. I don't think this is enough of a reason, though.
- foxglacier 1y agoIf you're deciding between using 8 bits or 16 bits, you might pick 16 because 8 is too small. But making the same decision between 9 and 18 bits could lead to picking 9 because it's good enough at the time. So no I don't think there would be half as many cases. They'd be different cases.
- oasisbob 1y agoThe IPv4 networking case is especially weird to think about because the early internet didn't use classless-addressing before CIDR. Thinking about the number of bits in the address is only one of the design parameters. The partitioning between network masks and host space is another design decision. The decision to reserve class D and class E space yet another. More room for hosts is good. More networks in the routing table is not. Okay, so if v4 addresses were composed of four 9-bit bytes instead of four 8-bit octets, how would the early classful networks shaken out? It doesn't do a lot of good if a class C network is still defined by the last byte.
- MangoToupe 1y agoMaybe if we worked with 7-bit bytes folks would be more grateful.
- xpe 1y agoFor those that don't get it, I'll explain. Imagine an alternative world that used 7-bit bytes. In that world, Pavel Panchekha wrote a blog post titled "We'd be Better Off with 8-bit Bytes". It was so popular that most people in that world look up to us, the 8-bit-byters. So to summarize, people that don't exist* are looking up to us now. * in our universe at least (see Tegmark's Level III Multiverse): https://space.mit.edu/home/tegmark/crazy.html https://space.mit.edu/home/tegmark/crazy.html or Wikipedia
- folsom 1y agoI don't know what if we ended up with a 27 bit address space? As far as ISPs competing on speeds in the mid 90s, for some reason it feels like historical retrospectives are always about ten years off.
- pavpanchekha 1y agoAuthor here, copied from another comment above. Actually I doubt we'd have picked 27-bit addresses. That's about 134M addresses; that's less than the US population (it's about the number of households today?) and Europe was also relevant when IPv4 was being designed. In any case, if we had chosen 27-bit addresses, we'd have hit exhaustion just a bit before the big telecom boom, a lucky coincidence meaning the consumer internet would largely require another transition anyway. Transitioning from 27-bit to I don't know 45-bit or 99-bit or whatever we'd choose next wouldn't be as hard as the IPv6 transition today.
- jayd16 1y agoI guess nibbles would be 3 bits and you'd 3 per byte?
- monocasa 1y agoOhh, and then we could write the digits in octal. Interestingly, the N64 internally had 9 bit bytes, just accesses from the CPU ignored one of the bits. This wasn't a parity bit, but instead a true extra data bit that was used by the GPU.
- ethan_smith 1y agoThe N64's Reality Display Processor actually used that 9th bit as a coverage mask for antialiasing, allowing per-pixel alpha blending without additional memory lookups.
- monocasa 1y agoAs well as extra bits in the Z buffer to give it a 15.3 fixed point format.
- kazinator 1y ago36 bit addresses would be better than 32, but I like being able to store a 64 bit double or pointer or integer in a word using NaN tagging (subject to the limitation that only 48 bits of the pointer are significant).
- Retr0id 1y agoAside from memory limits, one of the problems with 32-bit pointers is that ASLR is weakened as a security mitigation - there's simply fewer bits left to randomise. A 36-bit address space doesn't improve on this much. 64-bit pointers are pretty spacious and have "spare" bits for metadata (e.g. PAC, NaN-boxing). 72-bit pointers are even better I suppose, but their adoption would've come later.
- AlotOfReading 1y agoASLR has downsides as well. The address sanitizers have a shadow memory overhead that depends on the entropy in the pointer. If you have too much entropy, it becomes impossible for the runtime linker to map things correctly. Generally they'll just disable ASLR when they start, but it's one of the problems you'd have to solve to use them in production like ubsan even though that'd be extremely useful.
- kazinator 1y agoProblem is, not only did we have decades of C code that unnecessarily assumed 8/16/32, this all-the-world-is-a-VAX view is now baked into newer languages. C is good for portability to this kind of machine. You can have a 36 bit int (for instance), CHAR_BIT is defined as 9 and so on. With a little bit of extra reasoning, you can make the code fit different machines sizes so that you use all the available bits.
- pratyahava 1y agowas that assumption in C code really unnecessary? i suppose it made many things much easier.
- kazinator 1y agoIn my experience, highly portable C is cleaner and easier to understand and maintain than C which riddles abstract logic with dependencies on the specific parameters of the abstract machine. Sometimes the latter is a win, but not if that is your default modus operandi. Another issue is that machine-specific code that assumes compiler and machine characteristics often has outright undefined behavior, not making distinctions between "this type is guaranteed to be 32 bits" and "this type is guaranteed to wrap around to a negative value" or "if we shift this value 32 bits or more, we get zero so we are okay" and such. There are programmers who are not stupid like this, but those are the ones who will tend to reach for portable coding.
- pratyahava 1y agoyep, i remember when i tried coding for some atmega, i was wondering "how big are int and uint?" and wanted the types names to always include the size like uint8. but also there is char type, which should become char8 which looks even more crazy.
- kazinator 1y agoWould you want the main function to be: int32_t main(int32_t argc, char **argv)? How about struct tm? struct tm {$ int32_t tm_sec; /* Seconds (0-60) */$ int32_t tm_min; /* Minutes (0-59) */$ int32_t tm_hour; /* Hours (0-23) */$ int32_t tm_mday; /* Day of the month (1-31) */$ int32_t tm_mon; /* Month (0-11) */$ int32_t tm_year; /* Year - 1900 */$ int32_t tm_wday; /* Day of the week (0-6, Sunday = 0) */$ int32_t tm_yday; /* Day in the year (0-365, 1 Jan = 0) */$ int32_t tm_isdst; /* Daylight saving time */$ }; What for? Or do we "shrink wrap" every field to the smallest type? "uint8_t tm_hour"?
- TruffleLabs 1y agoPDP-8 has a 12-bit word size
- bawolff 1y ago> But in a world with 9-bit bytes IPv4 would have had 36-bit addresses, about 64 billion total. Or we would have had 27 bit addresses and ran into problems sooner.
- bigstrat2003 1y agoThat might've been better, actually. The author makes the mistake of "more time would've made this better", but we've had plenty of time to transition to IPv6. People simply don't because they are lazy and IPv4 works for them. More time wouldn't help that, any more than a procrastinating student benefits when the deadline for a paper gets extended. But on the other hand, if we had run out sooner, perhaps IPv4 wouldn't be as entrenched and people would've been more willing to switch. Maybe not, of course, but it's at least a possibility.
- dmitrygr 1y ago> simply don't because they are lazy and IPv4 works for them Or because IPv6 was not a simple "add more bits to address" but a much larger in-places-unwanted change.
- zamadatix 1y agoMost of the "unwanted" things in IPv6 aren't actually required by IPv6. Temporary addresses, most of the feature complexity in NDP, SLAAC, link-local addresses for anything but the underlying stuff that happens automatically, "no NAT, you must use PD", probably more I'm forgetting. Another large portion is things related to trying to be dual stack like concurrent resolutions/requests, various forms of tunneling, NAT64, and others. They're almost always deployed though because people end up liking the ideas. They don't want to configure VRRP for gateway redundancy, they don't want a DHCP server for clients to be able to connect, they want to be able to use link-local addresses for certain application use cases, they want the random addresses for increased privacy, they want to dual stack for compatibility, etc. For the people that don't care they see people deploying all of this and think "oh damn, that's nuts", not realizing you can still just deploy it almost exactly the same as IPv4 with longer addresses if that's all you want.
- smallstepforman 1y agoThe elephant in the room nobody talks about is silicon cost (wires, gates, multiplexirs, AND and OR gates etc). With a 4th lane, you may as well go straight to 16 bits to a byte.
- pratyahava 1y agoThis must be the real reason of using 8-bit. But then why did they make 9-bit machine instead of 16-bit?
- AlotOfReading 1y agoThe original meaning of byte was a variable number of bits to represent a character, joined into a larger word that reflected the machine's internal structure. The IBM STRETCH machines could change how many bits per character. This was originally only 1-6 bits [1] because they didn't see much need for 8 bit characters and it would have forced them to choose 64 bit words, when 60 bit words was faster and cheaper. A few months later they had a change of heart after considering how addressing interacted with memory paging [2] and added support for 8 bit bytes for futureproofing and 64 bit words, which became dominant with the 360. [1] https://web.archive.org/web/20170404160423/http://archive.computerhistory.org/resources/text/IBM/Stretch/pdfs/06-08/102632289.pdf https://web.archive.org/web/20170404160423/http://archive.co... [2] https://web.archive.org/web/20170404161611/http://archive.computerhistory.org/resources/text/IBM/Stretch/pdfs/06-08/102632292.pdf https://web.archive.org/web/20170404161611/http://archive.co...
- xpe 1y agoWould you lay out your logic (pun intended) a bit more? In what cases does doing from 8-bit bytes to 9-bit bytes result in something like a 2X penalty? One possibility would be bit-indexed addressing. For the 9-bit case, yes, such an index would need 4 bits. If one wanted to keep nice instruction set encoding nice and clean, that would result in an underutilized 4th bit. Coming up with a more complex encoding would cost silicon. What other cases are you thinking of?
- 1y ago
- SlowTao 1y agoCan you imagine the argument for 8bit bytes if we still lived in the original 6bit world of the 1950s? A big part of the move to 8bit systems was that it allowed expanded text systems with letter casing, punctuation and various ASCII stuff. We could move to the world of Fortran 36bit if really needed and solve all these problems while introducing a problem called Fortran.
- LegionMammal978 1y agoThere was already more than enough space for characters with 12-bit systems like the PDP-8. If anything, the convergence on 8-bit words just made it more efficient to use 7-bit codepages like ASCII.
- consp 1y agoAs the UTF encodings have shown you can put any encoding in any bitform if need be.
- duskwuff 1y agoNon-power-of-2 sizes are awkward from a hardware perspective. A lot of designs for e.g. optimized multipliers depend on the operands being divisible into halves; that doesn't work with units of 9 bits. It's also nice to be able to describe a bit position using a fixed number of bits (e.g. 0-7 in 3 bits, 0-31 in 5 bits, 0-63 in 6 bits), e.g. to represent a number of bitwise shift operations, or to select a bit from a byte; this also falls apart with 9, where you'd have to use four bits and have a bunch of invalid values.
- falcor84 1y agoWe just need 3 valued electronics
- percentcer 1y agoon, off, and the other thing
- tyingq 1y agohi-z is one choice. Though I don't know how well that does past a certain speed.
- duskwuff 1y agoIt works poorly at any speed. Hi-Z is an undriven signal, not a specific level, so voltage-driven logic like (C)MOS can't distinguish it from an input that's whatever that signal happens to be floating at. In current-driven logic like TTL or ECL, it's completely equivalent to a lack of current.
- tyingq 1y agoI wasn't pitching it as a solid commercial idea. Just that you can get (perhaps fiddly) three states out into the real world with something cheap that already exists. Like: https://idle-spark.blogspot.com/2015/04/low-cost-n-ary-dacs-using-digital-io.html https://idle-spark.blogspot.com/2015/04/low-cost-n-ary-dacs-...
- alphazard 1y agoWhen you stop to think about it, it really doesn't make sense to have memory addresses map to 8-bit values, instead of bits directly. Storage, memory, and CPUs all deal with larger blocks of bits, which have names like "pages" and "sectors" and "words" depending on the context. If accessing a bit is really accessing a larger block and throwing away most of it in every case, then the additional byte grouping isn't really helping much.
- SpaceNoodled 1y agoIt makes sense for the address to map to a value the same width as the data bus. A one-bit wide bus ... er, wire, now, I guess ... Could work just fine, but now we are extremely limited with the number of operations achievable, as well as the amount of addressable data: an eight-bit address can now only reference a maximum of 32 bytes of data, which is so small as to be effectively useless.
- alphazard 1y agoIf each memory address mapped to a CPU word sized value, that would make sense, and that is closer to the reality of instructions reading a word of memory at a time. Instead of using the CPU word size as the smallest addressable value, or the smallest possible value (a bit) as the smallest addressable value, we use a byte. It's an arbitrary grouping, and worse, it's rarely useful to think in terms of it. If you are optimizing access patterns, then you are thinking in terms of CPU words, cache line sizes, memory pages, and disk sectors. None of those are bytes.
- wvenable 1y agoThere was a time, however, where CPUs operated almost exclusively on on 8bit bytes (and had 8bit data buses). Everything else is merely the consequence of that.
- dboreham 1y agoQuick note that this isn't true. 8-bit CPUs are newer than 36-bit CPUs, and 8-bit bytes were established long before 8-bit CPUs came on the scene. Prior to the mid-70s most machines were word-addressed. The adoption of byte addressing (the subject of the grandparent) gained traction after C and similar languages became popular. C was developed on the pdp-11 which supported byte addressing and it provided a compatible memory model: any pointer value could be de-referenced to a byte. The VAX followed, also with byte addressing and by 1980 you couldn't sell a CPU that didn't support byte addressing (because C wouldn't work with it). 8-bit CPUs had nothing to do with any of this.
- m463 1y agoWe have already solved this problem many times. In clothing stores, numerical clothes sizes have steadily grown a little larger. The same make and model car/suv/pickup have steadily grown larger in stance. I think what is needed is to silently add 9-bit bytes, but don't tell anyone. also: https://imgs.xkcd.com/comics/standards_2x.png https://imgs.xkcd.com/comics/standards_2x.png
- nottorp 1y agoOf course, if that happens we'll get an article demanding 10-bit bytes. Got to stop somewhere.
- NelsonMinar 1y agoThis is ignoring the natural fact that we have 8 bit bytes because programmers have 8 fingers.
- classichasclass 1y agoNo, we still have 10. Real programmers think in octal. ;)
- mkl 1y agoMost have 10. That's the reason we use base 10 for numbers, even though 12 would make a lot of things easier: https://en.wikipedia.org/wiki/Duodecimal https://en.wikipedia.org/wiki/Duodecimal
- alserio 1y agoISO reserves programmers thumbs to LGTM on pull requests
- childintime 1y agoAvailable. As they suck on their thumbs.
- skort 1y ago> Thank you to GPT 4o and o4 for discussions, research, and drafting. Note to the author, put this up front, so I know that you did the bare minimum and I can safely ignore this article for the slop it is.
- js8 1y agoI have thought for fun about a little RISC microcomputer with 6-bit bytes, and 4-byte words (12 MiB of addressable RAM). I think 6-bit bytes would have been great at a point in history, and in something crazy fun like Minecraft. (It's actually interesting question, if we were to design early microprocessors with today's knowledge of HW methods, things like RISC, caches or pipelining, what would we do differently?)
- zokier 1y agoAnother interesting thought experiment would what if we went down to 6 bit bytes instead? Then the common values probably would be 24 and especially 48 bits (4 and 8 bytes), but 36 bit values might have appeared also in some places. In many ways 6 bit bytes would have had similar effect than 9 bit bytes; 18 and 36 bits would have been 3 and 6 bytes instead of 2 and 4 bytes. Notably with 6 bit bytes text encoding would have needed to be multibyte from the get-go, which might have been significant benefit (12 bit ASCII?)
- wmf 1y agoSome early mainframes used 6-bit characters which is why they didn't have lowercase.
- pavpanchekha 1y agoAuthor here. I agree that this would have a similar effect; we'd probably still end up with 36-bit or 48-bit IP addresses (though 30-bit would have been possible and bad). We'd probably end up with a transition from 24-bit to 48-bit addresses. 18-bit Unicode still seems likely. Not sure how big timestamps would end up being; 30-bit is possible and bad, but 48-bit seems more likely.
- deleted 1y ago[deleted]
- HappyPanacea 1y agoPanchekha is on a roll lately, I just read all of his recent posts a week ago. I really liked his AI vs Herbie series.
- pavpanchekha 1y agoThanks
- sedatk 1y agoOur capability to mispredict wouldn't have been different. We would have still picked the wrong size, and got stuck with scaling problems.
- labrador 1y agoAt the end: "Thank you to GPT 4o and o4 for discussions, research, and drafting." At first I thought that was a nice way to handle credit, but on further thought I wonder if this is necessary because the base line assumption is that everyone is using LLMs to help them write.
- xandrius 1y agoYeah, I don't remember ever thanking the spellchecker anything in the past. Maybe we are kinder to technology nowadays that we even credit it? Thank you to Android for mobile Internet connectivity, browsing, and typing.
- labrador 1y agoA counter point is that googling "thank you linux" turns up a lot of hits. "thank you linux for opening my eyes to a bigger world" is a typical comment.
- svachalek 1y agoAs soon as that's my baseline assumption, I think I'm done with the internet. I can get LLM slop on my own.
- labrador 1y agoI thought the article was well written. I'm assuming the author did most of the writing because it didn't sound like AI slop. I also assume he meant he uses AI to assist, not as the main driver.
- jacquesm 1y agoIt really wasn't well written. I contains factual errors that stand out like lighthouses showing the author had an idea about an article but doesn't actually know the material.
- 1y ago
- PaulHoule 1y agoI thought the PDP 10 had 6-bit bytes, or at least 6-bit characters https://en.wikipedia.org/wiki/Six-bit_character_code#DEC_SIXBIT_code https://en.wikipedia.org/wiki/Six-bit_character_code#DEC_SIX... Notably the PDP 8 had 12 bit words (2x6) and the PDP 10 had 36 bit words (6x6) Notably the PDP 10 had addressing modes where it could address a run of bits inside a word so it was adaptable to working with data from other systems. I've got some notes on a fantasy computer that has 48-bit words (fit inside a Javascript double!) and a mechanism like the PDP 10 where you can write "deep pointers" that have a bit offset and length that can even hang into the next word, with the length set to zero bits this could address UTF-8 character sequences. Think of a world where something like the PDP 10 inspired microcomputers, was used by people who used CJK characters and has a video system that would make the NeoGeo blush. Crazy I know.
- xdennis 1y agoThis is what happens when you write articles with AI (the article specifically mentions ChatGPT). The article says: > A number of 70s computing systems had nine-bit bytes, most prominently the PDP-10 This is false. If you ask ChatGPT "Was the PDP-10 a 9 bit computer?" it says "Yes, the PDP-10 used a 36-bit word size, and it treated characters as 9-bit bytes." But if you ask any other LLM or look it up on Wikipedia, you see that: > Some aspects of the instruction set are unusual, most notably the byte instructions, which operate on bit fields of any size from 1 to 36 bits inclusive, according to the general definition of a byte as a contiguous sequence of a fixed number of bits. -- https://en.wikipedia.org/wiki/PDP-10 https://en.wikipedia.org/wiki/PDP-10 So PDP-10 didn't have 9-bit bytes, but could support them. Characters were typically 6 bytes, but 7-bit and 9-bit characters were also sometimes used.
- vincent-manis 1y agoActually, the PDP-10 didn't have any byte size at all, it was a word-addressed machine. (An early attempt to implement C on this machine came a cropper because of this.) It did have a Load Byte and a Store Byte instruction, which allowed you to select the byte size. Common formats were Sixbit (self-explanatory), ASCII (5 7-bit bytes and an unused bit), and (more rarely, I think), 9-bit bytes. My first machines were the IBM 7044 (36-bit word) and the PDP-8 (12-bit word), and I must admit to a certain nostalgia for that style of machine (as well as the fact that a 36-bit word gives you some extra floating-point precision), but as others have pointed out, there are good reasons for power-of-2 byte and word sizes.
- nayuki 1y agoToday, we all agree that "byte" means 8 bits. But half a century ago, this was not so clear and the different hardware manufacturers were battling it out with different sized bytes. A reminder of that past history is that in Internet standards documents, the word "octet" is used to unambiguously refer to an 8-bit byte. Also, "octet" is the French word for byte, so a "gigaoctet (Go)" is a gigabyte (GB) in English. (Now, if only we could pin down the sizes of C/C++'s char/short/int/long/long-long integer types...)
- gnabgib 1y agoAn octet is unambiguously Latin for 8 of something; instruments, players, people, bytes, spider's legs, octopus' arms, molecules (see: octane). Octad/octade was unambiguously about 8 bit bytes, but fell out of popular usage.
- Tor3 1y agoThe term "octet" is still widely used in protocol descriptions and some other fields (source: All those interface specifications I have to read through my job)
- Dwedit 1y agoMany old 8-bit processors was basically 9-bit processors once you considered the carry flag.
- apt-apt-apt-apt 1y agoWe may have been stuck with slower, more expensive machines for 40+ years while computers that couldn't fully use the higher limits wasted time and energy.
- brudgers 1y agoI'm writing this on a four-year-old Macbook Pro and it only has 16 GB of RAM. Server-class machines would still need to address more memory than that, but they're usually running specialized software or virtualizing; databases and hypervisors are already tricky code and segmentation wouldn't be the end of the world. Because, I have a ten year old Dell laptop with 40GB of RAM, 16GB seems like an arbitrary limitation, an engineering compromise, or something like that. I don’t see how it is a result of 8 bit bytes because 64bits has a lot of address space. And because my laptop is running Windows 10 currently and ram Ubuntu before that, ordinary operating systems are sufficient. —- Also ECC RAM is 9 bits per byte.
- kyralis 1y ago"We've guessed wrong historically on data sizes, and if we had 9 bit bytes those guesses (if otherwise unchanged) would have been less wrong, so 9 bit bytes would be better!" is an extremely tenuous argument. Different decisions would have been made. We need to be better at estimating require sizes, not trying to trick ourselves into accomplishing that by slipping in an extra bit to our bytes.
- pavpanchekha 1y agoAuthor here. The argument was that by numerological coincidence, a couple of very important numbers (world population, written characters, seconds in an epoch, and plausible process memory usage) just happen to lie right near 2^16 / 2^32. I couldn't think of equally important numbers (for a computer) near ~260k or ~64B. We just got unlucky with the choice of 8-bit bytes.
- Nevermark 1y ago1 extra (literally) odd bit would require a lot of changes... What if instead of using single bytes, we used "doublebytes"? 8-bit software continues to work, while new 16-bit "doublebyte" software gets 256x the value capacity, instead of a meager 2x. Nobody will ever need more byte space than that! Without requiring any changes to CPU/GPU, RAM, SSD, Ethernet, WiFi ... Magic. :)
- LarMachinarum 1y agowhile none of the arguments of the article came even close to being convincing or to balancing out the disadvantages of a non-power-of-two orientation, there actually is one totally different argument/domain where the 9 bit per byte thing would hold true, that is: ECC bits in consumer devices (as opposed to just on servers): The fact that Intel managed to push their shitty market segmentation strategy of only even supporting ECC RAM on servers has rather nefarious and long-lasting consequences.
- shmerl 1y ago> It's 2025 and Github—Github!—doesn't support IPv6 Yeah, I wonder why. It's not IPv6's problem though, it's definitely Github's. Anyway, it's not a good example, since IPv6 is vastly wider than 9-bit variant of IPv4 would have been.
- crazygringo 1y agoIt's an interesting observation that 2^16 = 65K is a number that isn't quite big enough for things it's mostly big enough for, like characters. And that 2^32 = 4B is similarly awkwardly not quite big enough for global things related to numbers of people, or for second-based timestamps. But a 9th bit isn't going to solve those things either. The real problem is that powers-of-two-of-powers-of-two, where we jump from 256 to 65K to 4B to 18QN (quintillion), are just not fine-grained enough for efficient usage of space. It might be nice if we could also have 2^12=4K, 2^24=16M, and 2^48=281T as more supported integer bit lengths used for storage both in memory and on disk. But, is it really worth the effort? Maybe in databases? Obviously 16M colors has a long history, but that's another example where color banding in gradients makes it clear where that hasn't been quite enough either.
- montag 1y agoNah…We would have attempted to squeeze even bigger things into 18- and 36-bit address spaces that would have been equally short-sighted. But this is a tragedy of successes :)
- deleted 1y ago[deleted]
- protocolture 1y agoYeah uh, moving ISP's from IPv4 to IPv6 has been a headache, moving backwards to IPv4-9BIT would fuck things even harder.
- adrianmonk 1y ago> IPv4 would have had 36-bit addresses, about 64 billion total. That would still be enough right now, and even with continuing growth in India and Africa it would probably be enough for about a decade more. [ ... ] When exhaustion does set in, it would plausibly at a time where there's not a lot of growth left in penetration, population, or devices, and mild market mechanisms instead of NATs would be the solution. I think it's actually better to run out of IPv4 addresses before the world is covered! The later-adopting countries that can't get IPv4 addresses will just start with IPv6 from the beginning. This gives IPv6 more momentum. In big, expensive transitions, momentum is incredibly helpful because it eliminates that "is this transition even really happening?" collective self-doubt feeling. Individual members of the herd feel like the herd as a whole is moving, so they ought to move too. It also means that funds available for initial deployment get spent on IPv6 infrastructure, not IPv4. If you try to transition after deployment, you've got a system that mostly works already and you need to cough up more money to change it. That's a hard sell in a lot of cases.
- elcritch 1y agoAnd nothing like FOMO of developing markets not being able to access a product to drive VPs and CEOs to care about ensuring IPv6 support works with their products.
- pavpanchekha 1y agoAuthor here. My argument in the OP was that we maybe would never need to transition. With 36-bit addresses we'd probably get all the people and devices to fit. While there would still be early misallocation (hell, Ford and Mercedes still hold /8s) that could probably be corrected by buying/selling addresses without having to go to NATs and related. An even bigger address space might be required in some kind of buzzword bingo AI IoT VR world but 36 bits would be about enough even with the whole world online.
- qwerty2000 1y agoNot a very good argument. Yes, more bytes in situations where we’ve been constrained would have relieved the constraint… but it would eventually come. Even IP addresses… we don’t need an IP per person… IPv6 will be IPs for every device… multiple even… including an interplanetary network.
- xwolfi 1y ago[flagged]
- mhandley 1y agoIf we had 9-bit bytes and 36-bit words, then for the same hardware budget, we'd have 12.5% fewer bytes/words of memory. It seems likely that despite the examples in the article, in most cases we'd very likely not make use of the extra range as 8/32 is enough for most common cases. And so in all those cases where 8/32 is enough, the tradeoff isn't actually an advantage but instead is a disadvantage - 9/36 gives less addressable memory, with the upper bits generally unused.
- layer8 1y agoIt would also make Base64 a bit simpler (pun intended), at the cost of a little more overhead (50% instead of 33%).
- hinkley 1y agoWe likely wouldn’t use base64 at all in that case, but Base256. But also more of Europe would have fit in ASCII and Unicode would be a few years behind.
- layer8 1y agoThe point of Base64 is to represent binary data using a familiar character repertoire. At least for the latin-script world, any collection of 256 characters won’t be a familiar character repertoire.
- hinkley 1y agoWho told you that? Don’t talk to that person anymore. Base64 and uuencode before it are about transmitting binary data over systems that cannot handle binary. There are a bunch of systems in the early Internet that could only communicate 7 bits per byte, which is why uuencode uses only printable low ASCII characters. Has nothing to do with familiarity. Systems that supported eight bits per byte were referred to as “8 bit clean”, to distinguish them from legacy systems you might still have to support. PNG file format was specced in 1995, and it was still worried about 8 bit clean transmission. The first byte of the PNG magic number has the high bit set because they didn’t want the decoder to even have to bother with broken PNG files. > Base64 is also widely used for sending e-mail attachments, because SMTP – in its original form – was designed to transport 7-bit ASCII characters only. Encoding an attachment as Base64 before sending, and then decoding when received, assures older SMTP servers will not interfere with the attachment. I think it’s reasonable to assume that in a world with 9 bit bytes, someone may have chosen 8 bits for SMTP, or moved to 8 bit sooner. Which would give you at least Base128.
- theultdev 1y ago> a little more overhead (50% instead of 33%) a little?
- cwmoore 1y ago“It goes up to 1011!”
- ForOldHack 1y ago[flagged]
- tomhow 1y agoWe detached this comment from https://news.ycombinator.com/item?id=44818867 https://news.ycombinator.com/item?id=44818867 and marked it off topic.
- croemer 1y ago> Though you still see RFCs use "octet" Author seems to be unaware that octet is etymologically linked to 8.
- miiiiiike 1y agoThis is just an argument for longer roads and can kicking. He thanks ChatGPT for "discussions, research, and drafting". A real friend would have talked him out of posting this.
- pavpanchekha 1y ago[flagged]
- phyzome 1y agoWhat do you think they missed? What they said seems accurate: You're putting a lot of faith in ChatGPT that doesn't seem warranted.
- pavpanchekha 1y agoIt's not an argument for longer roads or can kicking? The thesis is stated at the top: with 9 bit bytes we'd avoid a couple of bad numerological coincidence (like that the number of Chinese characters is vague but plausibly around 50k, or the population of the earth is small integer billions). By avoiding the coincidences we'd avoid certain problems entirely. Unicode is growing but we're not discovering another Chinese! Not will world population ever hit 130B, at least as it looks now. If you think there's some equally bad coincidence go ahead and tell me but no one has yet. I think I do a good job of that in the post. (Also it's amazing you and maybe everyone else assume I know nothing except what ChatGPT told me? There are no ads on the website, it's got my name face and job on it, etc. I stand by what I wrote.)
- phyzome 1y agoPerhaps I misinterpreted your reply? It sounded like you were saying "read better" in response to the ChatGPT part, not the can-kicking part. ...but I do agree about the can-kicking. Why not 4, or 10? You'd have different pros and cons with each.
- klik99 1y agoThen people would be saying we’d be better off with 10-bits! Seriously though we can always do more with one more bit. That doesn’t mean we’d be better off. 8-bits is a nice symmetry with powers of two
- xwolfi 1y ago> Thank you to GPT 4o and o4 for discussions, research, and drafting. Yeah okay, this is completely pointless... so now we have to verify everything this guy published ?
- joshu 1y agosomeone suggested 10-bit bytes. this will not be enough. 11-bit bytes should be plenty, though
- sentrysapper 1y agoI will think about this post every time I hear the expression "it won't change things one bit".
- willvarfar 1y ago10-bit bytes would also be tempting, as 1024 is so close to 1000 and it would make bytes follow an orders of magnitude progression. It would just be so much easier on mental arithmetic too.
- billseng 1y agoWhat if the first bits told you how long the byte was, and you just kept reading the length of the byte until you got some terminating sequence, then that would be followed by some error correction for the length and terminating sequence with a couple more terminating sequences, one could be ignored since once might be corrupted, following by a variable length byte, with its own error correction and more error correction? It’s just so obvious y’all!
- ryukoposting 1y agoAnd if a foot were 13 inches, the numbers on your car's speedometer would be smaller. What is the point of this post?
- djcapelis 1y agoMost proposals for 9 bit bytes weren't for adopting 8 bits of data in a byte, they were to have 8 bits for data and 1 bit for something else, typically either error detection or differentiating between control/data. Very few folks argued for 9 bit bytes in the sense of having 9 bits of data per byte. 9 bit bytes never made significant headway because a 12.5% overhead cost for any of these alternatives is pretty wild. But there are folks and were folks then who thought it was worth debating and there certainly are advantages to it, especially if you look at use beyond memory storage. (i.e. closer to "Harvard" architecture separation between data / code and security implications around strict separation of control / data in applications like networking.) It's worth noting that SECDED ECC memory adds about a 20% overhead, though it can correct single bit flips whereas 9-bit bytes with a parity bit can only detect (but not correct) bit flips which makes it useful in theory but not very useful in practice.
- calebh 1y agoPerhaps the reason modern programs use so much memory vs what I remember from the Windows XP era is precisely because we went to 64 bits. Imagine how many pointers are used in the average program. When we switched over to 64 bits, the memory used by all those pointers instantly doubled. It's clear that 32 bits wasn't enough, but maybe some intermediate number between 32 and 64 would have added sufficient capacity without wasting a ton of extra space.
- frutiger 1y agoBecause of aligned reads any pointer size between 32-bit and 64-bit would end up using 64-bits anyway.
- IshKebab 1y agoWith 8-bit bytes, yes.
- zozbot234 1y ago> Imagine how many pointers are used in the average program. When we switched over to 64 bits, the memory used by all those pointers instantly doubled. This is a very real issue (not just on the Windows platform, either) but well-coded software can recover much of that space by using arena allocation and storing indexes instead of general pointers. It would also be nice if we could easily restrict the system allocator to staying within some arbitrary fraction of the program's virtual address space - then we could simply go back to 4-byte general pointers (provided that all library code was updated in due course to support this too) and not even need to mess with arenas. (We need this anyway to support programs that assume a 48-bit virtual address space on newer systems with 56-bit virtual addresses. Might as well deal with the 32-bit case too.)
- Tor3 1y agoSGI used three ABIs for their 64-bit computers.. O32, N32, N64. N32 was 64-bit except for pointers which were still 32 bits for exactly that reason - to avoid doubling the memory needed for storing pointers.
- 0xDEAFBEAD 1y ago9-bit bytes could easily make the common case ~12.5% slower everywhere. Most integers and chars fit in 8 bits no problem.
- rich_sasha 1y agoI often wondered why not make the word 10 bits, so that unsigned char is 0..1023 (basically 1,000). I guess hardware considerations make sense.
- Duanemclemore 1y agoIf we had gone the way of -1, 0, and 1 like some Soviet systems did this would be two "bits" Look I'm not a computer scientist, I admit this is naive. But for the thought experiment...
- usr1106 1y agoDidn't follow how Github not supporting IPv6 is caused by the "wrong" byte size. Wouldn't 36 bit IP adresses have made that a non-topic? The author seems to assume Github! is a leader. The masses in IT never have followed leading technology. How many Microsoft engineers do you need to change a light bulb? Zero, MS makes darkness an industry standard. Are Github actions the leading CI technology?
- necovek 1y agoI don't think at the time ASCII was being "upgraded" with localized 8-bit codepages, Greek would have had primacy over, say, Cyrillic. I wonder what came first, CP737 for Greek or CP855 and CP866 for Cyrillic.
- mparramon 1y agoThey call it the programmer's dozen, 9 bits for a byte
- gblargg 1y agoThis article's setup seems to be: we could go back and change bytes to be 9 bits, but make all the same decisions for sizes of things as we did, so that everything would be the same now except we'd have a little more room.
- ajuc 1y ago10 bit is even better according to all these criteria AND it fits 0-1000 into one byte which meshes really well with metric system (0 to 1 km in meters, 0 to 1 liter in ml, etc.) You could even do binary-encoded-metric-numbers that you can decode as needed one byte at a time - the first byte is tonnes, the second is kilograms, the third is grams, the 4th is milligrams, and you only lose 23 out of 1024 values at each level. Same (but without loses) with data sizes. 1st bit is gigabytes, 2nd is megabytes, 3rd is kilobytes, 4th is bytes. And of course at one point many computers used 40-bit floating point format which would fit nicely into our 4 bytes. 10-bit bytes would consist of two 5-bit nibbles, which you could use for two case-insensitive letters (for example Baudot Code was 5-bit). So you could still do hex-like 2-letter representation. Or you could send case-insensitive letters at 2 letters per byte. 40 bit could address 1 TB of memory (of 10-bit values - so much more than 1TB of 8-bit values). We could still be on 4-byte memory addressing to this day which would make all pointers 4-byte which would save us memory. And so on. But ultimately it always had to be power-of-two for cheaper hardware.
- danlitt 1y agoIf bytes weren't 8 bits, why would IPv4 addresses contain 4 bytes? Shouldn't they contain 3, or 9?
- 1GZ0 1y agoBut why stop there? based on your arguments 10-bit bytes would be even more better.
- childintime 1y agoI can think of only one flavor in favor of 9 bit bytes: variable length integers. The 9th bit would indicate there is more data to come. This would apply to instructions too. A homo iconic ISA, anyone?
- ordu 1y agoI used to think, how the history of computing and Internet would look like, if computers converged on 3-base system, with trits instead of bits, and trytes instead of bytes. If one tryte was 9 trites, it would have 3^3=19693 values. All the European characters and a lot of others can be encoded with this. There would be no need to invent char/int integer types in C (with the added mess of short, short short, long, and long long) int would be enough at the time. Maybe at the point when it would become necessary to add different integer types, C would choose a saner approach of stdint.h, and there would be no legacy code playing with legacy integer types? And 27 trites (or 3 trytes) is around 2^42.8 values, like 42 bits. It would be enough even now, I think.
- fnord77 1y agobinary is far easier to do in electronics (on or off)
- dpassens 1y agoWhat exactly is the argument here? Bigger numbers are bigger? But then we could also argue that by that logic, octets are better than 9-bit bytes because you need more bytes sooner and that gives you seven more bits over the additional one in 9-bit. > Thank you to GPT 4o and o4 for discussions, research, and drafting. That explains a lot.
- npn 1y agoThis is such a dumb article. All of those examples are nonsense. We can also have thousands of examples about how 6 bits is enough or 10 bits is just right.
- praptak 1y agoLisp implementors would love additional bits for tagging pointers more efficiently.
- blurbleblurble 1y agoIf 9 bits sounds nice wait until you hear about 16 bits
- mxfh 1y agoI just consider ourselves lucky, that we're not stuck with 6- or 7-bit bytes in ASCII-land and made it to Code page 437. https://en.wikipedia.org/wiki/Six-bit_character_code https://en.wikipedia.org/wiki/Six-bit_character_code
- account42 1y agoA measly factor 16 doesn't really make it worth having to deal with non-power of two sizes. You're also assuming that everything would have used the same number of bites when most sizes are chosen based on how much was needed at the time or in the foreseeable future - with 9 bit bytes that would just have meant that we're just going to run out earlier for different things than with 8 bit bytes. > IPv4: Everyone knows the story: IPv4 had 32-bit addresses, so about 4 billion total.44 Less due to various reserved subnets. That's not enough in a world with 8 billion humans, and that's lead to NATs, more active network middleware, and the impossibly glacial pace of IPv6 roll-out. It's 2025 and Github—Github!—doesn't support IPv6. But in a world with 9-bit bytes IPv4 would have had 36-bit addresses, about 64 billion total. That would still be enough right now, and even with continuing growth in India and Africa it would probably be enough for about a decade more. Only if you assume there is only one device per human, which is ridiculous. > Unicode: In our universe, there are 65 thousand 16-bit characters, which looked like maybe enough for all the world's languages, assuming you're really careful about which Chinese characters you let in.77 Known as CJK unification, a real design flaw in Unicode that we're stuck with. With 9-bit bytes we'd have 262 thousand 18-bit characters instead, which would totally be enough—there are only 155 thousand Unicode characters today, and that's with all the cat smileys and emojis we can dream of. UTF-9 would be thought of more as a compression format and largely sidelined by GZip. Which would be a lot worse than the current situation because most text like data only uses 8 bits per character. Text isn't just what humans type and includes tons of computer generated ASCII constructs. Not to mention that now it becomes an active process to upgrade ASCII data to Unicode, which would have the argument of increased size against it for many files and thus files and formats without Unicode support would have stuck around for much longer. UTF-8 might have been an accident of history in many ways but we really couldn't have wished for something better.
- hn92726819 1y agoThis article reminds me a lot of a conversation with an expert DBA recently. We were talking about performance improvements and he said that upgrading hardware is almost negligible. Upgrading from HDD to SSD gives you 10x boost. SSD to nvme maybe 4x. But optimizing queries and fixing indexes can give you 100x or 1000x improvements. Of course, every little bit counts, so we still did upgrade the hardware, but (at least in our area), almost all speed improvements came from code optimization. This article is talking about kicking the ipv4 can down down the road only 10 years and increasing process memory from 2G to 32G. Seems like such small fries when we could just double it and move on. If you brought the 2038 problem to Unix devs, I'm sure they would have said "thanks! We'll start with 64-bit" instead of "yes... Let's use an unaligned 36 bits so the problem is hidden slightly longer".
- stackedinserter 1y agoI worked on 29 bit machine and never actually noticed it, probably because it's the least of your problem when you use 1960's hardware lol
- whoomp12342 1y ago10-bt bytes seems a bit more logical as atleast its a recognizable base
- entaloneralie 1y agoLet's go ternary!
- vrajspandya1 1y agoHacker News needs, “potentially slop” button.
- infogulch 1y agoRather than extra bits, we may end up with ternary computers for AI. That's right, ternary is back with a new name: 1.58 bits