9 ms·
A bug fix in the 8086 microprocessor, revealed in the die's silicon
- MBCook 4y agoSo this was all to avoid having to re-layout everything and cut new rubylith, right? And at this point was all still done by hand? I suppose you’d have to re-test everything with either a new layout or this fix, so no real cost to save there?
- dboreham 4y agoYes. A hardware patch.
- dezgeg 4y agoAnd even still today this kind of manual hand-patching will be done for minor bug fixes (instead of 'recompiling' the layout from the RTL code). One motivator is if the fixes are simple enough, the modifications can be done such they only affect the top-most metal layer, so there is no need to remanufacture the masks for other layers, saving time and $$$.
- retrac 4y ago> And at this point was all still done by hand? Sort of. From what I've gathered, at this point Intel was still doing the actual layout with rubylith, yes. But in small modular sections. The sheets were then digitized and stitched into the final whole in software; there wasn't literally a giant 8086 rubylith sheet pieced together by hand, unlike just a few years before with the 8080. But the logic and circuits etc. were on paper, and there was no computer model capable of going from that to layout. The computerized mask was little more than an digital image. So a hardware patch it would have to be, unless you want to redo a lot. Soon, designers would start creating and editing such masks directly in CAD software. But those were just giant (for the time) images, really, with specialized editors. Significant software introspection and abstraction handling them came later. I don't think the modern approach of full synthesis from a specification, really came into use until the late 80s.
- marcosdumay 4y agoI remember the Pentium was marketed as "the first computer entirely designed on CAD", well into the 90's. But I'm not sure how real was that marketing message.
- eric__cartman 4y agoThat's a good advertisement for 486s because they were baller enough to pull that off.
- rasz 4y ago286 was an RTL model hand converted module by module to transistor/gate level schematic. Afair 386 was the first Intel CPU where they fully used synthesis (work out of UC Berkeley https://vcresearch.berkeley.edu/faculty/alberto-sangiovanni-vincentelli https://vcresearch.berkeley.edu/faculty/alberto-sangiovanni-... co-founder of a little company called Cadence) instead of manual routing. Everything went thru logic optimizers (multi-level logic synthesis) and will most likely be unrecognizable. I found a paper on this 'Coping with the Complexity of Microprocessor Design at Intel – A CAD History' https://www.researchgate.net/profile/Avinoam-Kolodny/publication/268005718_Coping_with_the_Complexity_of_Microprocessor_Design_at_Intel_-_A_CAD_History/links/54a13aa70cf256bf8bae6779/Coping-with-the-Complexity-of-Microprocessor-Design-at-Intel-A-CAD-History.pdf https://www.researchgate.net/profile/Avinoam-Kolodny/publica... >In the 80286 design, the blocks of RTL were manually translated into the schematic design of gates and transistors which were manually entered in the schematic capture system which generated netlists of the design. >Albert proposed to support the research at U.C. Berkeley, introduce the use of multi-level logic synthesis and automatic layout for the control logic of the 386, and to set up an internal group to implement the plan, albeit Alberto pointed out that multi-level synthesis had not been released even internally to other research groups in U.C. Berkeley. >Only the I/O ring, the data and address path, the microcode array and three large PLAs were not taken through the synthesis tool chain on the 386. While there were many early skeptics, the results spoke for themselves. With layout of standard cell blocks automatically generated, the layout and circuit designers could myopically focus on the highly optimized blocks like the datapath and I/O ring where their creativity could yield much greater impact >486 design: A fully automated translation from RTL to layout (we called it RLS: RTL to Layout Synthesis) No manual schematic design (direct synthesis of gate-level netlists from RTL, without graphical schematics of the circuits) Multi-level logic synthesis for the control functions Automated gate sizing and optimization Inclusion of parasitic elements estimation Full chip layout and floor planning tools
- _Microft 4y agoA Twitter thread by the author can be found here: https://twitter.com/kenshirriff/status/1596622754593259526 https://twitter.com/kenshirriff/status/1596622754593259526
- raphlinus 4y agoOr, if you prefer Mastodon: https://mastodon.online/@kenshirriff@oldbytes.space/109412325585692106 https://mastodon.online/@kenshirriff@oldbytes.space/10941232...
- kens 4y agoUsually if I post a Twitter thread, people on HN want a blog post instead. I'm not used to people wanting the Twitter thread :)
- jiggawatts 4y agoBlogs are too readable. I prefer the challenge of mentally stitching together individual sentences interspersed by ads.
- IIAOPSW 4y agoWell I've got great news for you (1/4)
- lazzlazzlazz 4y agoThe Twitter thread is better. You get the full context of other peoples' reactions, comments, references, etc. This is lost with blogs. It's the same reason why many people go to the Hacker News comments before clicking the article. :)
- Shared404 4y agoI like having the twitter/mastodon thread easily linked, but 10/10 times would rather see the article first.
- 4y ago
- techwiz137 4y agoAh yes, the infamous mov ss, pop ss that caused many a debuggers to fail and be detected.
- fnordpiglet 4y agoI sincerely wish I were this person. Reading this makes me feel I’ve fundamentally failed in my life decisions.
- _Microft 4y agoIt's never too late.
- fnordpiglet 4y agohttps://youtu.be/n3SsJdm8bMY https://youtu.be/n3SsJdm8bMY
- deleted 4y ago[deleted]
- kens 4y agoI'm not sure if I should be complimented or horrified by this.
- redanddead 4y agowhy not onboard him into the initiative?
- holoduke 4y agoI actually had some eye openers on how a cpu works in more detail after reading this article. Is there any good material for understanding cpu design for beginners?
- johannes1234321 4y agoCheckout Ben Eaters series on building a breadboard computer: https://youtube.com/playlist?list=PLowKtXNTBypGqImE405J2565dvjafglHU https://youtube.com/playlist?list=PLowKtXNTBypGqImE405J2565d... In the series he builds all the things more or less from the group up. And if that isn't enough he in another video series builds his breadboard VGA graphics card
- ajross 4y agoI love these so much. kens: were you able to identify the original bug? I understand the errata well enough, it implies that the processor state is indeterminate between the segment assignment and subsequent update to SP. But this isn't a 2/386 with segment selectors; on the 8086, SS and SP are just plain registers. At least conceptually, whenever you access memory the processor does an implicit shift/add with the relevant segment; there's no intermediate state. Unless there is? Was there a special optimization that pre-baked the segment offset into SP, maybe? Would love to hear about anything you saw.
- wbl 4y agoIt's not that the state is indeterminate it's that a programmer would have some serious trouble ensuring that it was possible to dump registers at that moment.
- ajross 4y agoNo, that would just be a software bug (and indeed, interrupt handlers absolutely can't just arbitrarily trust segment state on 8086 code where app code messes with it, for exactly that reason). The errata is a hardware error: the specified behavior of the CPU after a segment assignment is clear per the docs, but the behavior of the actual CPU is apparently different.
- mjw1007 4y agoThe 8086 didn't have a separate stack for interrupts. So, as I understand it, the problem isn't that code running in the interrupt handler might push data to a bad place; the problem is that the the processor itself would do so when pushing the return address before branching to the interrupt handler.
- cesarb 4y ago> the problem is that the the processor itself would do so when pushing the return address before branching to the interrupt handler. I want to emphasize this point: when handling an interrupt, the x86 processor itself pushes the return address and the flags into the stack. Other processor architectures store the return address and the processor state word into dedicated registers instead, and it's code running in the interrupt handler which saves the state in the stack; that would easily allow fixing the problem in software, instead of requiring it to be fixed in hardware (a simple software fix would be to use a dedicated save area or interrupt stack, separate from the normal stack; IIRC, it's that approach which later x86 processors followed, with a separate interrupt stack).
- userbinator 4y agoThe obvious workaround for this problem is to disable interrupts while you're changing the Stack Segment register, and then turn interrupts back on when you're done. This is the standard way to prevent interrupts from happening at a "bad time". The problem is that the 8086 (like most microprocessors) has a non-maskable interrupt (NMI), an interrupt for very important things that can't be disabled. Although it's unclear whether the very first revisions of the 8088 (not 8086) with this bug ended up in IBM PCs, since that would be a few years before its introduction, the original PC and successors have the ability to disable NMI in the external logic via an I/O port.
- _tom_ 4y agoThis was in the early IBM PCs. I know, because I remember I had to replace my 8088 when I got the 8087 floating point coprocessor. I don't recall exactly why this caused it to hit the interrupt bug, but it did.
- deleted 4y ago[deleted]
- anyfoo 4y agoMaybe because IBM connected the 8087's INT output to the 8086's NMI, so the 8087 became a source of NMIs (and with really trivial stuff like underflow/overflow, forced rounding etc., to boot). Even if you could disable the NMI with external circuitry, that workaround quickly becomes untenable, especially when it's not fully synchronous.
- _tom_ 4y agoThat sounds right. It's been a while.
- kaszanka 4y agoFor those who are curious, it's the highest bit in port 70h (this port is also used as an index register when accessing the CMOS/RTC): https://wiki.osdev.org/CMOS#Non-Maskable_Interrupts https://wiki.osdev.org/CMOS#Non-Maskable_Interrupts
- kklisura 4y ago> While reverse-engineering the 8086 from die photos, a particular circuit caught my eye because its physical layout on the die didn't match the surrounding circuitry. Is there a software that builds/reverses circuity from die photos or this is all manual work?
- speps 4y agoCheck out the Visual 6502 project for some gruesome details: http://www.visual6502.org/ http://www.visual6502.org/ The slides are a great summary: http://www.visual6502.org/docs/6502_in_action_14_web.pdf http://www.visual6502.org/docs/6502_in_action_14_web.pdf
- kens 4y agoCommercial reverse-engineering places probably have software. But in my case it's all manual.
- anyfoo 4y agoOnce again, absolutely amazing. Those are more details of a really interesting internal CPU bug than I could have ever hopes for. Ken, do you think in some future it might be feasible for a hobbyist (even if just a very advanced one like you) to do some sort of precise x-ray imaging that would obviate the need to destructively dismantle the chip? For a chip of that vintage, I mean. Obviously that's not an issue for 8086 or 6502, since there are more than plenty around. But if there were ever for example an engineering sample appearing, it would be incredibly interesting to know what might have changed. But if it's the only one, dissecting it could go very wrong and you lose both the chip and the insight it could have given.[1] Also in terms of footnotes, I always meant to ask: I think they make sense as footnotes, but unlike footnotes in a book or paper (or in this short comment), I cannot just let my eyes jump down and back up, which interrupts flow a little. I've seen at least one website having footnotes on the side, i.e. in the margin next to the text that they apply to. Maybe with a little JS or CSS to fully unveil then. Would that work? [1] Case in point, from that very errata, I don't know how rare 8086s with (C)1978 are, but it's conceivable they could be rare enough that dissolving them to compare the bugfix area isn't desirable.
- molticrystal 4y ago>some sort of precise x-ray imaging that would obviate the need to destructively dismantle the chip I don't know about a hobbyist, but are you talking about something along the lines of "Ptychographic X-ray Laminography"? [0] [1] [0] https://spectrum.ieee.org/xray-tech-lays-chip-secrets-bare https://spectrum.ieee.org/xray-tech-lays-chip-secrets-bare [1] https://www.nature.com/articles/s41928-019-0309-z https://www.nature.com/articles/s41928-019-0309-z
- anyfoo 4y agoI haven't ever looked into it itself, but what you just pasted seems like a somewhat promising answer indeed. Except for the synchrotron part I guess? Maybe?
- sbierwagen 4y ago>Except for the synchrotron part I guess? Maybe? They were using 6.2keV x-rays from the Swiss Light Source, a multi-billion Euro scientific installation. 6.2keV isn't especially energetic by x-ray standards, (a tungsten target x-ray tube will do ten times that) so either they needed monochromacy or the high flux you can only get from a building-sized synchrotron. Given that the paper says they poured 76 million Grays of ionizing radiation into an area of 90,000 cubic micrometers over the course of 60 hours suggests the latter. (A fatal dose of whole-body radiation to a human is about 5 Grays. This is not a tomography technique that will ever be applied to a living subject, though there is some interesting things no doubt being done right now to frozen bacteria or viruses.)
- ilyt 4y agoI'm more surprised that it already had microcode-equivalent, even if it was essentially hardcoded
- drmpeg 4y agoWhen I was at C-Cube Microsystems in the mid 90's, during the bring-up of a new chip they would test fixes with a FIB (Focused Ion Beam). Basically a direct edit of the silicon.
- mepian 4y agoIntel is still using FIB for silicon debugging.
- hota_mazi 4y agoI started reading this article and immediately thought "That's something Ken Shiriff could have written", but I didn't recognize the domain name. Lo and behold... looks like Ken just changed his domain name. Amazing work, as always, Ken.
- sgtnoodle 4y agoJust brainstorming a software workaround. It seems like you could abuse the division error interrupt, and manipulate the stack pointer from inside its ISR? I assume the division error interrupt can't be interrupted since it is interrupt 0?
- kens 4y agoThere's no stopping an NMI (non-maskable interrupt), even if you're in a division error interrupt.
- sgtnoodle 4y agoThat's some funny priority inversion! Reading up on it, the division error interrupt will be processed by the hardware first (pushed to the stack), but then be interrupted by NMI before the ISR runs.
- colejohnson66 4y ago“Interrupt 0” wasn’t its priority, but it’s index. IIRC, the 8086 doesn’t have the concept of interrupt priority. Upon encountering a #DE, the processor would look up interrupt 0 in the IVT (located at 0x00000 though 0x00003 in “real mode”).
- klelatti 4y agoFantastic work by Ken yet again. It's striking how many of the early microprocessors shipped with bugs. The 6502 and 8086 both did and it was an even bigger problem for processors like the Z8000 and 32016 where the bugs really helped to hinder their adoption. And the problem of bugs was a motivation for both the Berkeley RISC team and for Acorn ARM team when choosing RISC.
- adrian_b 4y agoAll the current processors are shipping with known bugs, including all Intel and AMD CPUs and all CPUs with ARM cores, also including even all the microcontrollers with which I have ever worked, regardless of manufacturer. The many bugs (usually from 20 to 100) of each CPU model are enumerated in errata documents with euphemistic names like "Specification Update" of Intel and "Revision Guide" of AMD. Some processor vendors have the stupid policy of providing the errata list only under NDA. Many of the known bugs have workarounds provided by microcode updates, a few are scheduled to be fixed in a later CPU revision, some affect only privileged code and it is expected that the operating system kernels will include workarounds that are described in the errata document, and many bugs have the resolution "Won't fix", because it is considered that they either affect things that are not essential for the correct execution of a program, e.g. the values of the performance counters or the speed of execution, or because it is considered that those bugs happen only in very peculiar circumstances that are unlikely to occur in any normal program. I recommend the reading of some of the "Specification Update" documents of Intel, to understand the difficulties of designing a bug-free CPU, even if much of the important information about the specific circumstances that cause the bugs is usually omitted.
- albert_e 4y agoFascinating! Are there any good documentary style videos that dive into some of such details of chips and microprocessor architectures?
- caf 4y agoI believe this one-instruction disabling of interrupts is called the 'interrupt shadow'.
- StayTrue 4y agoFantastic read. I’ll note for others this blog has an RSS feed (to which I’m now subscribed!).