5 ms·
Isn't why this problem even exits the exact opposite? Intel was losing on the mobile market and changed internal testing to iterate faster by cutting corners.
by spdy 9y ago
Isn't why this problem even exits the exact opposite?
Intel was losing on the mobile market and changed internal testing to iterate faster by cutting corners.
Found a quote:
"We need to move faster. Validation at Intel is taking much longer than it does for our competition. We need to do whatever we can to reduce those times… we can’t live forever in the shadow of the early 90’s FDIV bug, we need to move on. Our competition is moving much faster than we are".
- mannykannot 9y agoThat is a very interesting perspective, and as far as I know it is correct, though perhaps Intel's situation in the mobile market was exacerbated by complacency?
- leoc 9y agoWhere’s that quote from? ISTR reading it (or something very similar) as reported speech in a HN comment. Overall it’s a depressing story of predictable market failure as well as internal misbehavior at Intel, if true. Few buyers want to pay or wait for correctness until a sufficiently bad bug is sufficiently fresh in human memory. And if you do want to, it’s not as if you’re blessed with many convenient alternatives.
- deeth_starr_v 9y agoThe quote is from the link above (referencing an anonymous reddit comment).
- mtgx 9y agoI think that could also have been the "official reason". The same reason could have been used to give the NSA some legroom for instance, but tell everyone that's why they won't do so much verification in the future.
- HelloNurse 9y agoObvious hypothesis: first complacency leads to incompetence, then starting to cut corners has catastrophic consequences. The two problems are wonderfully complementary. As other comments suggest, there might be a third stage, completely forgetting how to design and validate chips properly.
- eximius 9y agoOr the system was designed poorly to begin with and now you're stuck with the design for backwards compatibility reasons.
- HelloNurse 9y agoI'd expect engineers that are aware of such serious bugs to spit on the grave of backwards compatibility. After all, the worst case impact would be smaller than the current emergency patches: rewriting small parts of operating systems with a variant for new fixed processors.
- fpoling 9y agoThis implies that ARM vendors do less validation. I guess ARM is just so much simpler that good enough validation can be done faster. So essentially this is payback time for Intel for keeping compatibility with older code and simpler to program architecture (stricter cache coherence etc.). It is like one can only have 2 of cheap, reliable, easy-to-program.
- pkaye 9y agoI'm sure ARM vendors have their own problems... it is just that they tend to be used in application specific products so the bugs are worked around. Having come from a firmware background I've worked are tons of ugly workarounds for serious bugs in validated hardware. Furthermore, I just a read an article (can't find the link) that certain ARM Cortex cores have this same issues as Intel.
- lmm 9y ago> This implies that ARM vendors do less validation. I guess ARM is just so much simpler that good enough validation can be done faster. More likely "good enough" is much lower because ARM users aren't finding the bugs. The workloads that find these bugs in Intel systems are: heavy compilation, heavy numeric computation, privilege escalation attackers on multi-user systems. Those use cases barely exist on ARM: who's running a compile farm on ARM, or doing scientific computation on an ARM cluster, or offering a public cloud running on ARM?
- kabdib 9y agoMan, you should see the errata for some ARM-based SOCs. It's amazing that they work at all. Vendor, in conversation: "We're pretty sure we can make the next version do cache coherency correctly." Me (paraphrased): "Don't let the door hit you in the ass on the way out." Management chain chooses them anyway, I spend the next year chasing down cache-related bugs. Fun.
- djsumdog 9y agoARM is such a shitstorm. At least the PC with UEFI is a standard. With every ARM device, you have to have a specialized kernel rom just for that device. There have been efforts made on things like PostmarketOS, but still in general, ARM isn't an architecture. It's random pins soldered to an SoC to make a single use pile of shit.
- madez 9y agoWhy is it an issue to need a different kernel image for each device? I don't see a problem as long as there is a simple mechanism to specify your device to generate the right image. It's already like that with coreboot/libreboot/librecore, and it worked just fine for me.
- kabdib 9y agoImagine that you are the person leading the team that's making an embedded system on an ARM SOC. It's not Linux, so you have your own boot code, drivers and so forth. It's not just a matter of "welp, get another kernel image." You're doing everything from the bare metal on up. (I should remark that there are good reasons for this effort. Such as: It boots in under 500ms, it's crazy efficient, doesn't use much RAM, and your company won't let you use anything with a GPL license for reasons that the lawyers are adamant about). So now you get to find all the places where the vendor documentation, sample code and so forth is wrong, or missing entirely, or telling the truth but about a different SOC. You find the race conditions, the timing problems, the magic tuning parameters that make things like the memory controller and the USB system actually work, the places where the cache system doesn't play well with various DMA controllers, the DMA engines that run wild and stomp memory at random, the I2C interfaces that randomly freeze or corrupt data . . . I could go on. It's fun, but nothing you learn is very transferrable (with the possible exception of mistrust of people at big silicon houses who slap together SOCs).