13 ms·
Uninitialized garbage on ia64 can be deadly (2004)
- vardump 10mo agoPretty surprising. So IA64 registers were 65 bit, with the extra bit describing whether the register contains garbage or not. If NaT (Not a Thing) is set, the register contents are invalid and that can cause "fun" things to happen... Not that this matters to anyone anymore. IA64 utterly failed long ago.
- msla 10mo agoIn case someone hasn't heard: https://en.wikipedia.org/wiki/Itanium https://en.wikipedia.org/wiki/Itanium > In 2019, Intel announced that new orders for Itanium would be accepted until January 30, 2020, and shipments would cease by July 29, 2021.[1] This took place on schedule.[9]
- ashleyn 10mo agoThere are modern VLIW architectures. I think Groq uses one. The lessons on what works and what doesn't are worth learning from history.
- addaon 10mo agoA more everyday example is the Hexagon DSP ISA in Qualcomm chips. Four-wide VLIW + SMT.
- bri3d 10mo agoVLIW works for workloads where the compiler can somewhat accurately predict what will be resident in cache. It’s used everywhere in DSP, was common in GPU for awhile, and is present in lots of niche accelerators. It’s a dead end for situations where cache residency is not predictable, like any kind of multitenant general purpose workload.
- vardump 10mo agoI meant narrowly only about IA64. There is sure some lessons learned value.
- 0dyl 10mo agoThe new TI C2000 F29 series of microcontrollers are VLIW
- msla 10mo agoIA64 was EPIC, which, itself, was a "lessons learned" VLIW design, in that it had things like stop bits to explicitly demarcate dependency boundaries so instructions from multiple words could be combined on future hardware with more parallelism, and speculative execution and loads, which, well, see the article on how the speculative loads were a mixed blessing. https://en.wikipedia.org/wiki/Explicitly_parallel_instruction_computing https://en.wikipedia.org/wiki/Explicitly_parallel_instructio...
- kragen 10mo agoIt matters to people designing new hardware and maybe new virtual machine instruction sets.
- nottorp 10mo agoOr to people caring about their software working on more than just Chrome. ... oh wait, on more than x86(64).
- ronsor 10mo agoYet another reason IA64 was a design disaster. VLIW architectures still live on in GPUs and special purpose (parallel) processors, where these sorts of constraints are more reasonable.
- nneonneo 10mo agoI mean, there is a reason why these sorts of constructs are UB, even if they work on popular architectures. The problems aren’t unique to IA64, either; the better solution is to be aware that UB means UB and to avoid it studiously. (Unfortunately, that’s also hard to do in C).
- awesome_dude 10mo agoThe bigger problem is that a user cannot avoid an application where someone was writing code with UB, unless they both have the source code, and expertise in understanding it.
- eru 10mo agoIsn't that a general problem?
- loeg 10mo agoIt's a very weird architecture to have these NAT states representable in registers but not main memory. Register spilling is a common requirement!
- mwkaufma 10mo agoI assume they were stored in an out-of-band mask word
- amluto 10mo agoHah, this is IA-64. It has special hardware support for register spills, and you can search for “NaT bits” here: https://portal.cs.umbc.edu/help/architecture/aig.pdf https://portal.cs.umbc.edu/help/architecture/aig.pdf to discover at least two magical registers to hold up to 127 spilled registers worth of NaT bits. So they tried. The NaT bits are truly bizarre and I’m really not convinced they worked well. I’m not sure what happens to bits that don’t fit in those magic registers. And it’s definitely a mistake to have registers where the register’s value cannot be reliably represented in the common in-memory form of the register. x87 FPU’s 80-bit registers that are usually stored in 64-bit words in memory are another example.
- deleted 10mo ago[deleted]
- Joker_vD 10mo agoRaymond Chen has a whole "Introduction to IA-64" series of posts on his blog, by the way. It's such an unconventional ISA that I am baffled that Intel seriously thought they would've been able to persuade anyone to switch to it from x86: it's very poorly suited for general-purpose computations. Number crunching, sure, but anything more freeform, and you stare at the specs and wonder how the hell the designers supposed this thing to be programmed and used.
- yongjik 10mo agoWell, they did persuade HP to ditch their own homegrown PA-RISC architecture and jump on board with Itanium, so there's that. I wonder how much that decision contributed to the eventual demise of HP's high performance server division ...
- classichasclass 10mo agoA lot, I think. PA-RISC had a lot going for it, high performance, solid ISA, even some low-end consumer grade parts (not to the same degree as PowerPC but certainly more so than, say, SPARC). It could have gone much farther than it did. Not that HP was the only one to lose their minds over Itanic (SGI in particular), but I thought they were the ones who walked away from the most.
- pjc50 10mo agoAm I right in thinking that the old PA-Semi team was bought by Apple, and are substantially responsible for the success of the M-series parts?
- scrlk 10mo agoAcquiring P.A. Semi got them Dan Dobberpuhl and Jim Keller, which laid a good design foundation. However, IMO, I'd lean towards these as the decisive factors today: 1) Apple's financial firepower allowing them to book out SOTA process nodes 2) Apple being less cost-sensitive in their designs vs. Qualcomm or Intel. Since Apple sells devices, they can justify 'expensive' decisions like massive caches that require significantly more die area.
- nayuki 10mo ago> The ia64 is a very demanding architecture. In tomorrow’s entry, I’ll talk about some other ways the ia64 will make you pay the penalty when you take shortcuts in your code and manage to skate by on the comparatively error-forgiving i386. https://devblogs.microsoft.com/oldnewthing/20040120-00/?p=40993 https://devblogs.microsoft.com/oldnewthing/20040120-00/?p=40... "ia64 – misdeclaring near and far data" https://devblogs.microsoft.com/oldnewthing/2004/01 https://devblogs.microsoft.com/oldnewthing/2004/01
- andikleen2 10mo agoEarly x86-64 Linux had a similar problem. The x86-64 ABI uses registers for the first 6 arguments. To support variable number of arguments (like printf) requires passing the number of arguments in an extra register (RAX), so that the callee can save the registers to memory for va_arg() and friends. Doing this for every call is too expensive, so it's only done when the prototype is marked as stdarg. Now the initial gcc implemented this saving to memory with a kind of duffs device, with a computed jump into a block of register saving instructions to only save the needed registers. There was no boundary check, so if the no argument register (RAX) was not initialized correctly it would jump randomly based on the junk, and cause very confusing bug reports. This bit quite some software which didn't use correct prototypes, calling stdarg functions without indicating that in the prototype. On 32bit code which didn't use register arguments this wasn't a problem. Later compiler versions switched to saving all registers unconditionally.
- veltas 10mo agoIn the SysV ABI for AMD64 the AL register is used to pass an upper bound on the number of vector registers used, is this related to what you're talking about?
- jcalvinowens 10mo agoAt least they made the stack grow in the right direction! Well, half of it, anyway...