6 ms·
I think 128 bit computers will come around eventually, despite it having been declared that 64 bit is "enough". Some pressures may come from: * Memory addressi
by bArray 5y ago
I think 128 bit computers will come around eventually, despite it having been declared that 64 bit is "enough". Some pressures may come from:
* Memory addressing - As the article suggests, addressing large amount of address space. Not just RAM, but disk control on the bit/byte level (something solid state drives may enable and new filesystems may take advantage of). There may also be applications where you want an etabyte disk as low-speed RAM.
* Multi-byte processing - accelerating instructions like AVX have shown the power of processing multiple bytes at a time. One can imagine that wider registers would accelerate these processes and would allow for multi-byte processes to happen in parallel.
* Gaming/simulation - We have seen quite a few examples where physics in games and simulations have broken down due to the inaccuracy of `double` for large values. I believe Minecraft physics for example used to become extremely unstable near the world border.
* Hashing - With `int32_t` and large amounts of data, you will see a lot of collisions. `int64_t` is lesser, but still likely. `int128_t` is rarer but still possible. `int256_t` (`long long` on a 128 bit processor) would be highly unlikely. Being able to compare hashes in just a few clock cycles would be awesome.
* Custom instructions - When programs can define a custom instruction to speed-up computation, 128 bits, or 16 bytes, could even be enough to contain the custom instruction and the payload.
These are just things I've noticed. I imagine there are others too. The prediction of 2043 is still quite realistic, I wouldn't be surprised if we beat it.
I was quite disappointed to see many Linux distros give up on 32 bit support because it was too much effort to support. It probably points towards some crappy code that is highly dependant on the platform.
- anamax 5y agoMulti-issue 64bit machines can compare 256 bits in "just a few cycles". (You do 2-4 compares per cycle and then combine the results.)
- zarzavat 5y agoIn my opinion none of these are good reasons for 128-bit computers, as they can be easily addressed by using 2x64 bit integers and software arithmetic, or bigints. The issue with 32-bit specifically is that the maximum value (2 billion for a signed int) is really small and hits all kinds of practical limits: there are more humans than that! Whereas 2^63 is big enough for almost all practical purposes. There were costs to the 64 bit transition in terms of higher memory usage so I don’t think we’ll see another transition because the costs at 128-bit would outweigh the benefits.
- hajile 5y ago> Memory addressing Current 64-bit CPUs generally don't go past 52-bits of memory addressing. That's 4.5 Petabytes and they still have 12 doublings to go. It is also possible to have memory addressing wider than your CPU. 8-bit chips almost universally did this. Even the P4 had 36 bits on a 32-bit CPU. > Multi-byte processing You can have wider vector load instructions without affecting the main CPU width > Gaming/simulation Legit point. IBM POWER has 128-bit IEEE floats and decimal hardware coprocessors, but I believe the rest of the CPU is still 64-bits wide. The easiest solution here is to allow those 256/512 vector units to do 128-bit floats in addition to 64-bit floats. > Hashing If you have a hashtable that needs 128-bits due to collisions, it is a MASSIVE table. It definitely won't be fitting into cache and probably not into RAM. A few extra instructions (or even a hundred extra instructions) to do the math with two 64-bit segments will still be nothing compared to the thousands of cycles reaching out to RAM or the millions/billions of cycles reaching out to the disk.
- gnufx 5y ago> Legit point. IBM POWER has 128-bit IEEE floats and decimal hardware coprocessors, but I believe the rest of the CPU is still 64-bits wide. Yes (POWER9+, I think), but also 128-bit "IBM" format. Transitioning to an IEEE ABI for ppc64le GNU/Linux has caused some pain in GCC and elsewhere recently. (I don't know how software long double is typically compiled and how fast the results are.)
- gizmo686 5y ago> Memory addressing - As the article suggests, addressing large amount of address space. Not just RAM, but disk control on the bit/byte level (something solid state drives may enable and new filesystems may take advantage of). There may also be applications where you want an etabyte disk as low-speed RAM. Think bigger. We could go all the way to a global, routable, address space. Such a system would need to be sparse to avoid terrible fragmentation. We could also burn a bunch of address space on segmentation. Why spend instructions on checked array accesses when you could just put your 10 element array between a terabyte worth of unmapped virtual address space. In practice, I suspect we will move to more heterogeneous register sizing, with increasingly large registers being used for computation while pointers stay at 64 bits or smaller (maybe with fatter pointers in niche applications, like a multi-machine runtime embedding the machine address in the upper bits)
- jcranmer 5y agoWe already have 128-bit computers if you're willing to count 128-bit vector units. (The internal data path of most CPUs is 128-bit these days). So the basic problem with 128-bit comes down to one thing: memory. There is a physical limitation on the amount of memory you can have ready to access; if you've paid attention, there has been basically no growth in the size of caches in the past decade. Doubling the memory size of your basic constituents means you cut the effective size of that cache in half, and effectively using that cache is the biggest barrier to performance. If you look at prior sizes, it's clear that 16-bits is way too small-that's 65,536 entries, which is easy to overflow in many cases. 32-bits gives you about 4 billion entries, which is generally sufficient for the vast majority of cases, but can overflow (it only takes ~1 second for a CPU to count to 4 billion). Overflowing 64-bits in a counter would take that same CPU a century to do, and it should be rapidly clear that there's not much that's going to use it. Consider your use case of address spaces. Memory can't get all that much smaller; we're starting to come up on laws of physics, so there's only a factor of 1,000× or so smaller that we can get. We're barely at the threshold right now of exhausting the existing 48-bit address spaces (or 256TiB of memory), and even then, that requires things like memory-mapping entire disk drives to eat up that much memory. Apply that factor of 1,000 shrinkage, and you still don't cross the 16EiB threshold of a 64-bit address space. The need to address that much memory just isn't in the cards, and if it is, the memory is likely to be so heavily non-uniform that you'd probably see a development towards segmented address spaces again anyways to avoid needing 128-bit pointers generally. The one use case that I consider the most likely is adding hardware support for quad-precision floating point numbers, since it already exists in several libraries as a soft float operation, it's well-understood, the datapaths can generally already handle it (a 128-bit vector FMA unit can relatively easily be extended to support a 128-bit scalar floating-point FMA anyways), and there's a small need for it.
- catlifeonmars 5y agoDumb question: why stop at 128? Why not go directly to 512, or even 1024 bits?
- Banana699 5y agoThere's cost in the CPU because - Every register must be usable as a pointer, so if you make your addresses (i.e. pointers) too big for no good reason you're wasting silicon. - A solution to the above is picking a subset of registers that can address memory and making them big and leaving others as is, but this complicates the architecture and makes it ugly and complex - Another solution is to alias names so that registers are accessible either as monolithic 128/512/1024 blobs or as smaller subslices (A32_1 is the least significant 32 bits of the 128-bit register A, A32_2 is the next 32 bit slice and so on). This lessens the waste because when you don't need a 128-bit register you have 4 32 bit register instead and there's no such thing as too many registers. x86 does something vaguely similar but it's super ugly for different reasons, I don't see any problems with an architecture designed like this from the start, someone better at Computer Architecture might. - All functional units in the CPU will have to be the size of the largest register N, this would be a waste if they don't have the ability to function instead as N/n parallel n-bit units instead, and designing this into the architecture might be complex and entailing a lot of decisions. (e.g. If a 128-bit adder is functioning as parallel 4 32-bit adders, where do the 4 separate overflow bits go?)
- R0b0t1 5y agoModern 128 bit designs are held back by trouble with heat dissipation and signal routing. Basically you need 2-8x the transistors depending on operation, and you need 4x the space or more to route it.
- hulitu 5y agoAnd a company like DEC to invest in research.
- Stevvo 5y agoThe problems with games are not due to the inaccuracy of FP64 'double'; they are all caused by FP32 'single'.