Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
brucedawson
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
16 ms
·
61.
▲
by
brucedawson
4y ago
Those tactics could work. They would be a bit expensive however, and it would be a shame to have all users paying the performance penalty because of a few bad pieces of software. And, this would not have caught the two errors in the assembl
62.
▲
by
brucedawson
4y ago
I think most assemblers know what functions are. There are usually directives to indicate this. They help with emitting symbols (you need the function name to be emitted or no-one can call it), stack unwind information, and other informatio
63.
▲
by
brucedawson
4y ago
It appears that the question of what is valid on a moved-from object is tricky. Here is one discussion: https://stackoverflow.com/questions/7027523/what-can-i-do-wi... FWIW, here is the move operator for the type
64.
▲
by
brucedawson
4y ago
It would be easy to have an "auto" clobber list option. The inline assembler would note which registers were touched (mostly trivial, a few instructions have implicit destinations) and would add them all to the clobber list. In ab
65.
▲
by
brucedawson
4y ago
Both bugs were programming errors in assembly language files. One was inline assembly that was missing entries from a clobber list, the other was an assembly function that lacked invocations of the macros that were supposed to be used to pr
66.
▲
by
brucedawson
4y ago
The compiler is making assumptions (which it is supposed to make) but nobody is enforcing the assumptions. The only player who could reasonable enforce the assumptions would be the compiler, in a special checking mode. I am not aware of a c
67.
▲
by
brucedawson
4y ago
In fact the fix to the WebRTC bug was to adjust the clobber list, thus telling the compiler to save the registers. Why the compiler didn't notice that registers were being used that weren't on the clobber list is unclear to me. Yo
68.
▲
by
brucedawson
4y ago
You could write an arbitrarily complex analyzer to try to find violations but I suspect that the halting problem means that you can never be sure you've found all errors. I think that either crude heuristics (found two bugs!) or UBSan
69.
▲
by
brucedawson
4y ago
I was wondering if anybody was going to point that out. It did occur to me that the CHECK was not technically valid due to that exact concern, but given that we control the compiler and the C++ library implementation and given that it'
70.
▲
by
brucedawson
4y ago
IIRC we do block DLLs that aren't signed by either Google or Microsoft in some of our processes. In other processes we can't because third-party DLLs are needed for shell extensions (utility processes) or accessibility (browser pr
71.
▲
by
brucedawson
4y ago
The webrtc fix was thematically similar in that the programmer declared what registers were trashed and then the compiler knows which registers need to be saved. I'm not sure why the compiler doesn't notice when registers are used
72.
▲
by
brucedawson
4y ago
Exactly. My understanding of the conventions and macros in those source files is that you declare what registers you will be trashing, and then the registers are saved/restored as required by that platform. On Linux it would be a NOP,
73.
▲
by
brucedawson
4y ago
That would be possible, but we'd have to install the software, then guess which binary was the culprit, and then have some way of finding the function boundaries. My crude analysis technique required on having symbols for chrome.dll to
74.
▲
by
brucedawson
4y ago
Correctness bugs are just one type of bug. Performance bugs are another type. The designation of what should be counted as a bug is less clear for performance bugs but if some operation is slower than it could/should be, and if this sl
75.
▲
by
brucedawson
4y ago
I develop Windows client software and I need to debug and profile that software, so I need to work on Windows. And, the development tools on Windows are excellent so I'm happy to do that. So, I want bugs like this fixed. Also, Linux ne
76.
▲
by
brucedawson
5y ago
Trivia: the Xbox 360 CPU design used 10 FO4, meaning that it could get less done per clock cycle than the Pentium 4. In order words, the Xbox 360 CPU was trying to beat the Pentium 4 at the clock-speed game (fewer stages per clock allows fo
77.
▲
by
brucedawson
5y ago
Sorry, I've never seen a non-NDA document with full details. Some information was released but it was a bit spotty.
78.
▲
by
brucedawson
5y ago
The material was all proprietary/NDA and I don't think it's ever been shared publicly, sorry. The main point I wanted to make with how I became the CPU expert was not actually the "reading multiple times" in order t
79.
▲
by
brucedawson
5y ago
2004 was a bit early for a Netflix binge. Plus, the power was out, so there goes that plan. The documentation that I read showed the pipelines in great deal, but they were badly presented such that the flows of data and time were obscured.
80.
▲
by
brucedawson
5y ago
Reading from L2 requires sending down a request and then waiting for the response. This makes me realize that the actual speed is 5.5 mm in 0.625 ns because it's a two cycle delay in each direction. I'll update the article. But, t
81.
▲
by
brucedawson
5y ago
Having more logical/architectural registers is great except for a few costs: 1) More bits to encode register numbers in instructions. Doubling the number of logical registers costs another two or three bits depending on how many regist
82.
▲
by
brucedawson
5y ago
If x87 and MMX registers don't support renaming then that means that they can't support OOO and speculative execution of these instructions. This is possible but seems unlikely to me. That is, even though not a lot of x87/MMX
83.
▲
by
brucedawson
5y ago
Register renaming is implemented in hardware. Because it is used on every instruction it is on the critical path and is probably hand-optimized. Here is some more reading on this topic: https://en.wikipedia.org/wiki/Reg
84.
▲
by
brucedawson
5y ago
Yep, exactly. The rule of thumb is that whenever a register is written to a new register mapping is created. For instruction #1 eax might be assigned to physical register 103. It maintains that identity for instruction #2. For instruction #
85.
▲
by
brucedawson
5y ago
Adding more logical registers is compatibility breaking. And, since you have to encode the register specifier in the instruction it means larger instructions (hence the compatibility breaking) which makes reading and decoding instructions s
86.
▲
by
brucedawson
5y ago
Ouch. Zombie handles are so much worse. Now we're talking serious memory, and since it isn't attributed to the process that leaks the handles it is easy to not realize what is happening. More details: as I explain in the first lin
87.
▲
by
brucedawson
5y ago
That would _probably_ be safe, but you have to be sure that the xdcbt instructions in the "special" function are far enough into the function that speculative execution can never reach there. Pipelines are way deeper than most peo
88.
▲
by
brucedawson
5y ago
If Intel ships a CPU with a bug in it then that is an expensive mistake. If they produce a bunch of CPUs with a bug (escape to silicon) that can also be an expensive mistake. That said: 1) I deal with CPU bugs pretty regularly on Chrome. So
89.
▲
by
brucedawson
5y ago
Making the instruction not speculatable would indeed be a hardware change, which there was not time for. So that was not an option. And, let's say they did that. All other loads/prefetches are done in the early stages of the pipel
90.
▲
by
brucedawson
5y ago
The question of how to properly implement an instruction like xdcbt is interesting. Undoing the damage would be both tricky and expensive. Only doing the L2-skipping when the instruction is executed (as opposed to speculatively executed) wo
More ›