5 ms·
I was hired by a bookie ;) I have a somewhat related story. Circa 1995 or so, I was working on software compiled on a 486DX computer. Back then the computers
by emcrazyone 11y ago
I was hired by a bookie ;)
I have a somewhat related story. Circa 1995 or so, I was working on software compiled on a 486DX computer. Back then the computers had L2 cache in SDRAM chips you plugged into the motherboard.
I was moonlighting with a few guys to develop software and each night myself and two friends met up in office we were renting. There, on one of those fold out tables, we toiled away each night writing code.
The code we worked on was for off shore gambling (aka a bookie). Back in the 90s you could walk into a convenient store and pick up an magazine; usually autotrader magazine. On the back was an 800 number you could call to place bets on football, hokey, or baseball. (Called HOBs bets in the industry).
Anyway, because of the nature of the software and the people we dealt with, things were a bit hairy at times. The bookie we worked for paid us generously but, at the same time, expected perfection (i.e. software that just worked).
One evening, as usual, the three of us met to work toward the next release of the software. The next release of the software was to include a feature for boxing bets. Mike Tyson was about to be release from jail and the bookie was anticipating various future business in taking boxing bets.
So this one evening we were doing a software build to release as beta software. I did the software build and ran a battery of tests we normally run and all worked fine. To transfer the software to the offshore network we used a dial up connection.
I pushed the software out which my bookie contact would run and make sure all the logic was correct.
A few days go by and we get a phone call that the software is not functioning. Sometimes it crashes (which was rare) or sometimes just strange things would happen such as strange blinking characters on the screen.
I collected enough information to try and replicate the steps necessary. You can imagine my customer wasn't too happy. So the three of us basically stopped working on the code to track down the bug. We spent a week looking for the bug we were sure was in the code somewhere. Back then we didn't really use a code repository but we did have a diff of the beta release and previous release.
Pouring over the code changes we just could not find the problem. It even got to the point where I would recompile the beta version, run it, and could actually get the code to crash. Comparing to earlier release it would not crash.
It was by chance that my fellow coder, sitting right next to me, started running similar tests on his machine while I was off making a food run. It was sort of normal for us to group up around my workstation when something needed a collective look.
Long story short it turned out that one of the L2 cache chips was causing my compiles to become corrupt in just the right way to cause the code to be not as we typed it in.
Anyway, thought I'd share. Although not a malicious act, it was one of the worst things to track down. Fast forward to my current life... I work in the embedded field and recently solved a bit-flip problem in one of the products my company produces that relies on NAND memory. That little incident from my past certainly has served me.... ;)
- jacquesm 11y agoThat's an ugly one. And one that could easily happen today.
- i336_ 11y agoThat kind of thing is often exploited in a security context. See: http://dinaburg.org/bitsquatting.html http://dinaburg.org/bitsquatting.html
- perlgeek 11y agoI'm curious, how did you manage to pin it down to an L2 cache in particular? (As opposed to general memory corruption, for example).
- emcrazyone 11y agoah good question. It was very, very difficult to pin down. We knew our compiler was the same on both computers and compiler flags were all the same. The two machines were purchased at the same time and we knew everything about them was the same. We used something called a Pharlap DOS extender so our software could use beyond the 1MB boundary. It was in fiddling with that memory extender that I began to suspect a memory failure. Changing it's parameters eventually got me to fairly repeatable way to get the issue to show up not only in my app but in 3rd party software began failing too. Also, while we mostly did our work from a DOS prompt and a DOS editor (called qc) we would sometimes run Windows (Windows 3.11) which had it's own way of accessing Expanded vs. Extended memory through a DOS driver. So by swapping between the Pharlap driver and the Microsoft config.sys driver did I begin to suspect memory failure. In other words, the machine became more noticeably more unstable. I don't recall the exact name but I began running something like memcheck on the computers. Basically a memory checker and although it did not reveal the problem in its entirety, the memchecker would crash on my machine and not any of the others. These were computers running just DOS with Novell Netware network drivers for DOS/Netwware only. I reasoned that memory was failing before the memory walker/checker was getting corrupt. When I swapped main memory with another computer and the problem didn't follow the memory, my only other option was to pull out the L2 cache chips. When I put them in the computer next to mine, and finally saw the problem show up there I knew. My friend owned the computer company that sold us the computers and he later was able to test and validate my findings.