Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
pbsd
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
15 ms
·
121.
▲
by
pbsd
8y ago
Most of the eSTREAM finalists ( http://www.ecrypt.eu.org/stream/index.html ) are not block-based. I think only Salsa20 matches that description. More generally, I think the distinction between a "block-based stream
122.
▲
by
pbsd
8y ago
CNL also added another AES unit, so you can now dispatch aesenc and its ilk to ports 0 and 1.
123.
▲
by
pbsd
8y ago
OCB2 is not a block cipher; it is an authenticated encryption scheme built on top of one. The scary thing here is not so much the error in the proof, which does not have many repercussions beyond OCB2, but that it went 14 years without bein
124.
▲
by
pbsd
8y ago
There is the case of LZCNT and BSR. On processors that do not support the former, LZCNT is interpreted as (REP) BSR, but the instructions have different behavior that could result in silent failures.
125.
▲
by
pbsd
8y ago
I will continue to mark the security of PCG as "unwilling to quantify", which is strictly below "better than nothing". Debating the virtues of the latter is therefore pointless.
126.
▲
by
pbsd
8y ago
"Not breakable in fewer than O(2^(n/2)) operations" is an imprecise security claim; "challenging" means nothing. "Challenging" may not even mean "computationally hard"---breaking a truncated LCG
127.
▲
by
pbsd
8y ago
"Challenging" is inherently unfalsifiable, and gives you very little information about what security you're getting. You can come up with a better-than-bruteforce attack, only to be met with "see? I was right all along,
128.
▲
by
pbsd
8y ago
An AEAD is consuming n bytes of input per cycle and producing at least n bytes of output, that is, the output ciphertext plus tag. The definition of authenticated encryption is, essentially, indistinguishability of the ciphertext plus M
129.
▲
by
pbsd
8y ago
Agner's instruction listings [1] and InstLatX64 [2] are good sources for a variety of chips. The 24 cycle figure above can be seen at https://github.com/InstLatx64/InstLatx64/blob/ee13abcfb1e2bb... [1]
130.
▲
by
pbsd
8y ago
It's still doing it, but now it uses `lahf` + `sahf` instead to (re)store the flags. These are better than pushf + popf for sure, but they cannot be used in general code because some early x86_64 chips forgot to implement them.
131.
▲
by
pbsd
8y ago
Where are you getting that from? `pushfd` is OK-ish, but `popfd` is microcoded, requires 9 uops, and can only be issued once every ~20 cycles. That's not what I would call well-pipelined. You could probably achieve the same result (ass
132.
▲
by
pbsd
8y ago
That is not the case. On a Skylake chip, a pushfd + popfd roundtrip costs 24 cycles.
133.
▲
by
pbsd
8y ago
The code in the blog post does not match what is actually benchmarked. The reason `u256_full_mul` is so much faster than the inline assembly version is that it omits the upper part of the result, due to how the benchmark is done (cf. https
134.
▲
by
pbsd
9y ago
To be pedantic, Keccak does have a 5 to 5 bit S-box. But like 3-Way, NOEKEON, Serpent, etc, it's disguised---bitsliced---as a short sequence of boolean operations.
135.
▲
by
pbsd
9y ago
This theory is cool, but I don't think it works, all things considered. PDEP and PEXT should have the same unfused behavior as SHLX, since they also do not change any flags, but they _do_ fuse. BEXTR should (or could) fuse, but doesn&#
136.
▲
by
pbsd
9y ago
Renaming is unrelated to my guess about the flags. The point is that there's a limit to how many inputs a fused uop can have, 3, and the flags register may become one input too many to be able to fuse the uops. For example, inc [
137.
▲
by
pbsd
9y ago
According to Agner when the uops go back to to the reorder buffer to get retired they are still treated as fused. However, if you look at the descriptions of the events [1, 2] a retired fused uop is counted as 2 retired uops. And yes, if yo
138.
▲
by
pbsd
9y ago
As far as I understand, with micro uop fusion the number of retired uops should be higher than the issued uops.
139.
▲
by
pbsd
9y ago
> The point is that the three instruction format and the single instruction format you would think would decode both to three micro-ops (or something like that), but instead the single instruction format can use a single fused micro op.
140.
▲
by
pbsd
9y ago
That seems fine on its own, but it somewhat undermines the paragraph with "but they have all followed the same basic mechanism", seeing as these earlier attacks relying on caches etc did not follow this mechanism. Changing "a
141.
▲
by
pbsd
9y ago
That's a much more defensible claim---one I have no issue with---but it was easy to misread the post and believe it was claiming something more general.
142.
▲
by
pbsd
9y ago
I'm gonna go and take issue with the claim that you were the first to come up with microarchitectural attacks, and that their story begins in 2004: - Dan Page published [1] in 2002, describing an attack on DES exploiting cache timings.
143.
▲
by
pbsd
9y ago
C++ is not allowed to do the small buffer optimization for std::vector<T>. You'll find a number of small_vector<T> classes from third parties, though. The reason is that std::swap(v1, v2) must never invalidate existing iter
144.
▲
by
pbsd
9y ago
I later realized that padding to 32 bytes didn't make much of a difference, it is the number of vacuous NOP opcodes generated that matters. So it can be simplified to mov rcx, 400000000-1 up: nop nop nop add [
145.
▲
by
pbsd
9y ago
> I presume you are seeing the instructions showing up as IDQ_MITE_UOPS? No, the legacy decoder is virtually unused. This loop was designed to take advantage of the rules of the DSB. If you have the LSD enabled I don't quite know wh
146.
▲
by
pbsd
9y ago
We are observing different effects. I suggest you apply the latest microcode update, seeing that you have LSD_UOPS being dispatched; the latest update disabled the LSD altogether, due to the SKL150 CPU errata. The LSD was also permanently d
147.
▲
by
pbsd
9y ago
The call shifts the loop from being limited at the backend level (mostly uops not being retired by having to wait for memory) to being limited at the frontend level. The core event to look for here is `IDQ_UOPS_NOT_DELIVERED.CORE`, which te
148.
▲
by
pbsd
9y ago
Being offline does not ensure security against unverified plaintext. Imagine a colossal implementation fuckup (e.g., using strncmp to verify tags, ala Nintendo), or perhaps a hardware glitch flipping the verification bit to always return tr
149.
▲
by
pbsd
9y ago
Kevin Igoe has been gone from the CFRG since mid-2015.
150.
▲
by
pbsd
9y ago
Probably not, I don't know. There's a better paper from a couple of years ago [1], which even manage to include code [2]. One of the attacks on NORX [3] was essentially exploiting a bigger-than-expected invariant subspace, it migh
More ›