6 ms·
> The threat resides in the chips’ data memory-dependent prefetcher, a hardware optimization that predicts the memory addresses of data that running code is lik
by midtake 3y ago
> The threat resides in the chips’ data memory-dependent prefetcher, a hardware optimization that predicts the memory addresses of data that running code is likely to access in the near future.
Are we nearing any sort of consensus that any form of speculation is bad? Is there a fundamentally secure way to do it?
- SV_BubbleTime 3y agoIsn’t the issue more akin to use after free? If the instructions and memory were wiped on prediction path failure wouldn’t that help?
- black_puppydog 3y agonope. the side effects (which can be access times in case of non-failure, or voltage changes, or or or...) would still happen
- pjc50 3y agoAbsolutely critical for performance, though. If there's a way out of this it might have to be better virtualization.
- bhawks 3y agoFor cryptographic applications yes. That is why people have spent significant effort to implement constant time algorithms to replace standard math and bitwise operations. At the hardware level any optimizations that change performance characteristics locally (how long the crypto operation directly takes) or non locally (in this case the secrets leak via observation of cache timings in the attacker's untrusted code) are unsafe. Intel DMPs already have a flag to turn off the same behavior that was exploited on the M1/M2. Which may suggest that the risk of this type of optimization was understood previously. Mixing crypto operations with general purpose computation and memory accesses is a fragile balance. Where possible try utilizing HSMs, yubikeys, secure enclaves - any specialized hardware that has been hardened to protect key material.
- amelius 3y ago> For cryptographic applications yes. Why only cryptographic applications? What if I'm writing a very sensitive e-mail, for instance?
- bhawks 3y agoCheck out https://www.qubes-os.org/ https://www.qubes-os.org/ for an operating system that tries to put as many layers of defense as possible between an end user's applications.
- tsimionescu 3y agoFor this type of attack to work, the algorithm being run needs to be very well understood, and the runtime of the algorithm needs to depend almost entirely on the secret key. In contrast, the timing of virtually any email operation is not dependent on the contents of the email, other than the size. That is, whether you wrote "my password is hunter2" or "my password is passwor", the timing of any operation running on this email will be identical.
- eru 3y ago> In contrast, the timing of virtually any email operation is not dependent on the contents of the email, other than the size. What about spell checkers etc? Or even just whatever runs to figure out where to break the lines?
- tsimionescu 3y agoPerhaps those could be attacked. It's possible though that it's not feasible, that the possible inputs leading to a certain timing signature are just too many to get any data out of it. Consider that those programs are not making any effort whatsoever to run in constant time, and yet no one has shown any timing attack against them. OpenSSL has taken great pains to have constant execution time, and yet subtle processor features like this still introduce enough time differences to recover the keys.
- datadeft 3y agoIt is bad, it is required for performance reasons. The questions is what could be the solution going forward, which is going to be a huge change anyway. I do not see a way out of this with our current architectures.
- crest 3y agoNeither block ciphers, nor stream ciphers, nor common public key algorithms (RSA, Ed25519) need or even profit from this. They just need fast access to the register-register math, maybe loop sequentially through all members of a fixed sized array a fixed number of times. The only thing those implementing such algorithms would probably like having is a few kiB of safe to access scratchpad memory for code and data. On entry to the crypto code copy the code and data there, enable a constant time mode for compute instructions and run the algorithm at full speed without worrying.
- PhilipRoman 3y agoMy personal opinion is that we should solve it the opposite way - don't run untrusted code in the first place (with rare exceptions like dedicating an entire cpu core and a region of memory to a virtual machine, etc). Speculation is one of many side channel attacks, who knows what kind of crazy RF-based exploits are out there. AFAIK we still haven't fully solved rowhammer. I think for "normal" users the main risk is JavaScript, which can (kind of) be mitigated in software without affecting the rest of the system, so no one really cares about these attacks. But the fundamental abstraction leak between physics and programming will always be there.
- bhawks 3y agoSo you browse the Internet with JavaScript turned off? You're a bigger person than I am ;). The risk here is that there are more individuals with the skills to take this type of attack and bring it to a browser near you. One apps data is another apps code.
- PhilipRoman 3y agoIndeed, the scary thing is that there is no theoretical limit to how sophisticated a side channel attack could be. Imagine all the timing data that could in theory be gathered from html layout engines and css, even without javascript, just by resource loading times. I would like to salute my shitty ISP for keeping me safe from timing attacks using their unreliable network infrastructure.
- robin_reala 3y agoThis attack is now why browsers segment caching into a combination of requesting domain and asset URL, rather than just caching the asset on its own. It slows down for example Google Fonts, but means that a site can’t check to see that you’ve visited a competitor by timing an asset load from their site to see whether it’s in the cache.
- pmontra 3y agoI use uMatrix and allow first party JS. When some sites break I open the matrix look at what they would like to load and allow one more origin and reload. An example: chatgpt works by allowing JS from the first party domain, *azure.com and oaistatic.com, which looks like something from OpenAI. It would like to load JS from three other domains but it works even if I don't allow them, so there is no need to let that code run.
- actualwitch 3y agoShould we make a petition for apple to make lockdown-like mode that disables speculative execution? I'll sign up for that.
- bawolff 3y agoFrom the gofetch website, apple has already done this for m3 chips.
- amelius 3y agoPerhaps we should do the crypto in constant time, and run all other applications using homomorphic encryption?
- tsimionescu 3y agoFHE is unusably slow even for the simplest operations, and there is no reason to be sure it will ever be fast enough for any normal computing.
- layer8 3y agoDisabling all hardware optimizations becomes an option long before homomorphic encryption becomes an option, performance-wise.
- mort96 3y agoThis isn't speculative execution. EDIT: The downvotes make no sense. What this bug has in common with Spectre is that it has to do with cache timing. But in Spectre, the cache is affected by speculative execution; with "GoFetch", it's the pre-fetcher pre-fetching things which look like memory addresses. Pre-fetching is not speculative execution.
- bawolff 3y agoUnless the comment was edited, the person you are responded to did not use the phrase "speculative execution"
- mort96 3y agoYou're right, it says speculation. I read it as speculative execution, I have never ever heard the term "speculation" applied to prefetching... But if they did mean to include pre-fetching in "speculation" then I retract my comment
- deleted 3y ago[deleted]
- crest 3y agoSpeculation is required to get even close to the single-thread throughput expected of any modern CPU for anything worth running on a general purpose CPU. The problem is that there is no formal specified model to reason over side-channels not even just for timing side-channels. Most ISAs doesn't specify the time it takes execute instructions. Lets assume I multiply two 64bit numbers. The CPU could just do it the same way every time and the worst-case has 4 cycles latency. It may also track if one of the factors is zero and dynamically replace the multiplication with a zeroing idiom that "executes" in 0 cycles when the scheduler learns that that either input is zero as an extreme example. Less radical it could track if the upper halves of registers are zero to fast-path smaller multiplications (e.g. 32bit x 32bit -> 64bit) and shave off a cycle. IIRC some PowerPC chips did that, but for the second argument only. The ISA allowed it. A realistic example are CPUs with data-dependent latency shift/rotate instructions. What do you do if an ISA doesn't specify if shift/rotate is constant time, but every implementation of it so far did it in constant time? Do you slowly emulate it out of paranoia that a future implementation may have variable latency? An other real-world example of this would be FPUs that have higher latency for denormalised numbers its just not relevant to (most) cryptographic algorithms. How the fuck are you supposed to build anything secure, useful, and fast enough from that?