Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
universal_sinc
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
universal_sinc
3y ago
This article highly underestimates the value of keeping 128b vector performance high. Most code doesn't get recompiled or compiled with the appropriate flags. There is significant overhead involved in supporting 1x512b operations, 2x25
2.
▲
by
universal_sinc
3y ago
The idea is to write a C++ model that that produces cycle accurate outputs of the branch predictor, core pipeline, queues, memory latency, cache hierarchy, prefetch behaviour, etc. Transistor level accuracy isn't needed as long as the
3.
▲
by
universal_sinc
3y ago
Absolutely! Chip designers have a several tools to do this. First, they create detailed software models (usually in C++) of their chips to estimate performance as closely as they can before laying out a single transitory. These models can r
4.
▲
by
universal_sinc
4y ago
Even 0.1ns is way slow. A modern silicon cmos gate will switch under 10ps, which is how we can fit 25+ gates in a single cycle at >3GHz. Everyone should remember that cpu frequency is not the same as the frequency a single gate can switc
5.
▲
Arm AArch64 Adds Memcpy() Instructions
(community.arm.com)
158 points
by
universal_sinc
5y ago
|
94 comments
6.
▲
by
universal_sinc
6y ago
Just so everyone is aware of the scales involved in modern circuits, an entire CPU Core like in Snapdragon 865 might only take up an area of ~3mm2, measuring just ~1.75mm across. 3mm gets you across the CPU and back again. We think in nanom
7.
▲
by
universal_sinc
6y ago
It's not as bad as you think. From a high-level: Modern Synthesis tools turn your RTL code (which is coded in an HDL or Hardware Description Language) into gates, and then map them to a library of "Standard Cells". These foun
8.
▲
by
universal_sinc
6y ago
The importance of AMD's use of chiplets should not be understated. Especially in the server space, it allows them to achieve far better yields on a monster L3 cache than Intel or any other competitor. Along with the modernized Zen UArc
9.
▲
by
universal_sinc
6y ago
They way modern CPU design works is that the development team maintains a cycle-accurate software model of the CPU. Changes can be made to the pipeline and memory system, those changes simulated against a real workload, and results reported