Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
eigenform
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
eigenform
1y ago
(Sorry for the self-plug but) I also wrote a bit about the behavior of PREFETCH recently in case anyone is interested in this sort of thing. See this example (for Linux on AMD): https://github.com/eigenform/perfect
32.
▲
by
eigenform
1y ago
Similar idea I think: organics depositing iron
33.
▲
by
eigenform
1y ago
That post[^1] linked by saagarjha above is talking about the case where the typed allocator (plus the layout of kernel memory, and whatever constraints on pointer arithmetic in the kernel) makes Spectre less useful. MTE itself isn't re
34.
▲
by
eigenform
1y ago
I think part of the argument is that doing a micro-op cache is not exactly cutting down on your power/area budget. (But then again, do the AMD e-cores have uop caches?)
35.
▲
by
eigenform
1y ago
Don't know this for certain, but I always assumed that x86 implementations get away with this by predecoding cachelines. If you're going to do prefetching in parallel and decoupled from everything else, might as well move part of
36.
▲
by
eigenform
1y ago
> [...] imagine that while you are loading 16 or 32 bytes from instruction cache, you need to predict the address of next loaded chunk in the same cycle, before you even see what you got from cache. Yeah, you [ideally] want to predict t
37.
▲
by
eigenform
1y ago
In general, probably co-design with software. Apple is in a position where they design microprocessors that are only going to be running MacOS/iOS.
38.
▲
by
eigenform
1y ago
Yep, you typically don't see it because we learned that it's easier to just assume that the default prediction is "not-taken." AFAICT if you're hinting that a branch is biased-taken, really the only thing you might
39.
▲
by
eigenform
1y ago
Think you're referring to the idea that "my compiler can know that some branch is always/never taken" and turn it into an unconditional control-flow instruction (either "always jump here", or "always con
40.
▲
by
eigenform
1y ago
I'm all for giving programmers a way to flush state, and maybe this is just a matter of taste, but I wouldn't characterize this as "taking care of the problem once and for all" unless there's a [magic?] way to recov
41.
▲
by
eigenform
1y ago
Also weird because pipelining is very literally "insert memories to cut some bigger process into parts that can occur in parallel with one another"
42.
▲
Exploiting Retbleed in the Real World
(bughunters.google.com)
3 points
by
eigenform
1y ago
|
0 comments
43.
▲
by
eigenform
1y ago
[Not a language designer or anything, but] I dunno, keeping some distinction between Python and the DSL seems useful, and probably makes it simpler to reason about. Too much syntactic sugar, and I can imagine it becomes easier to confuse
44.
▲
by
eigenform
1y ago
This reminds me: has anyone ever figured out why Zen 3 was missing memory renaming, but it came back in Zen 4 and Zen 5?
45.
▲
by
eigenform
1y ago
Even more obvious if you think about the case of hardware-managed caches! The ISA typically exposes some simple cache control instructions (and I guess non-temporal loads/stores?), but apart from that, the actual choice of storage loca
46.
▲
RUNLTS: Register-Value-Aware Predictor Utilizing Nested Large Tables [pdf]
(ericrotenberg.wordpress.ncsu.edu)
4 points
by
eigenform
1y ago
|
0 comments
47.
▲
by
eigenform
1y ago
I'm not sure this sentence [from the paper] makes a lot of sense. The only thing non-standard is the use of Chisel (and then probably CIRCT to lower it into Verilog) - if you're actually taping these out, you're still feeding
48.
▲
Adventures in the Design of Ultra-Precision Machine Tools [video]
(youtube.com)
2 points
by
eigenform
1y ago
|
0 comments
49.
▲
by
eigenform
1y ago
Great read! Some boiled-down takeaways: - Predictor updates may be deferred until sometime after a branch retires. Makes sense, otherwise I guess you'd expect that branches would take longer to retire! - Dispatch-serializing instruct
50.
▲
by
eigenform
2y ago
The abstract mentions: > Here we present a full-waveform seismic tomographic model So presumably, you would actually do the seismic tomography (if they haven't already). Instead of radar, you use the waves from earthquakes!
51.
▲
by
eigenform
2y ago
> are there any similarities? Don't know about the format, but if you look thru old ITJ articles[^1], it seems like the "direct access" interface for reading out different memories exists on older Pentium parts too. Presum
52.
▲
by
eigenform
2y ago
I'm not qualified to say anything about this stuff, but I get the impression this is just a tradeoff: - Collect bits from many "points" in space, but a single "point" in time - Collect bits from a single "point
53.
▲
by
eigenform
2y ago
> Sorry, but you did not read carefully that good article and you did not read the AMD documentation and the Intel documentation. I think you are the one who hasn't read documentation or tested the behavior of Zen cores. Read lite
54.
▲
by
eigenform
2y ago
AFAIK you have to think about how many different 512b paths are being driven when this happens, like each cycle in the steady-state case is simultaneously (in the case where you can do two vfmadd132ps per cycle): - Capturing 2x512b from t
55.
▲
by
eigenform
2y ago
Mhm, agreed :)
56.
▲
by
eigenform
2y ago
Shawn mentions lack of observed ground deformation, but maybe he means near Kolumbo? There's a note[^1] on the EMSC site that has some figures for deformation at stations on Santorini, and it seems like it has been deflating while all
57.
▲
by
eigenform
2y ago
I'm contending that, like any good tool, there is a context where it is useful, and a context where it is not (and that we are at a stage where everything looks suspiciously like a nail).
58.
▲
by
eigenform
2y ago
It depends on your tolerance for error. When you have a machine that can only infer rules for reasoning from inputs [which are, more often than not, encoded in a very roundabout way within a language which is very ambiguous, like English]
59.
▲
by
eigenform
2y ago
They're trying to validate that you're using a trusted version of AGESA. This is probably intentional, the AMD bulletin[^1] mentions this (ie. for Milan): > Minimum MilanPI_1.0.0.F is required to allow for hot-loading future m
60.
▲
by
eigenform
2y ago
> for x86-64 CPUs it is very unlikely that a load value predictor can be worthwhile I think you're making a good point about immediate encodings probably making ARM code more amenable to LVP, but I'm not sure I totally buy this
More ›