4 ms·
These days CPUs are so complex and have so many interdependencies that the best way to simulate them is simply to run them! In most real code the high throughp
by barbegal 2y ago
These days CPUs are so complex and have so many interdependencies that the best way to simulate them is simply to run them!
In most real code the high throughput of these sorts of operations means that something else is the limiting factor. And if multiplier throughput is limiting performance then you should be using SIMD or a GPU.
- sroussey 2y ago> These days CPUs are so complex and have so many interdependencies that the best way to simulate them is simply to run them! True. You can imagine how difficult it is for the hardware engineer designing and testing these things before production!
- alain94040 2y agoVery true. To paraphrase a saying, CPU amateurs argue about micro-benchmarks on HN, the pros simulate real code.
- pizlonator 2y agoOr even just A:B test real code.
- devit 2y agoThe amateurs usually run benchmarks (because they can't reason about it as they lack the relevant knowledge) and believe they got a useful result on some aspect when in the fact the benchmark usually depends on other arbitrary random factors (e.g. maybe they think they are measuring FMA throughput, but are in fact measuring whether the compiler autovectorizes or whether it fuses multiply and adds automatically). A pro would generally only run benchmarks if it's the only way to find out (or if it's easy), but isn't going to trust it unless there's a good explanation for the effects, or unless they actually just want to compare two very specific configurations rather than coming up with a general finding.
- LegionMammal978 2y agoThen again, 'reasoning about it' can easily go awry if your knowledge doesn't get updated alongside the CPU architectures. I've seen people confidently say all sorts of stuff about optimization that hasn't been relevant since the early 2000s. Or in the other direction, some people treat modern compilers/CPUs like they can perform magic, so that it makes no difference what you shovel into them (e.g., "this OOP language has a compiler that always knows when to store values on the stack"). Benchmarks can help dispel some of the most egregious myths, even if they are easy to misuse.
- Joker_vD 2y ago> the best way to simulate them is simply to run them! And it's quite sad because when you are faced with choosing between two ways to express something in the code, you can't predict how fast one or another option will run. You need to actually run both, preferrably in an environment close to the prod, and under similar load, to get accurate idea which one is more performant. And the worst thing is, you most likely can't extract any useful general principle out of it, because any small perturbation in the problem will result in a code that is very similar yet has completely different latency/throughput characteristics. The only saving grace is that modern computers are really incredibly fast, so layers upon layers of suboptimal code result in applications that mostly perform okay, with maybe some places where they perform egregiously slow.
- mpweiher 2y ago> you can't predict how fast one or another option will run "The best way to predict the future is to invent it" -- Alan Kay > You need to actually run both Always! If you're not measuring, you're not doing performance optimization. And if you think CPUs are bad: try benchmarking I/O. Operating System (n) -- Mechanism designed specifically to prevent any meaningful performance measurement (every performance engineer ever) If it's measurable and repeatable, it's not meaningful. If it's meaningful, it's not measurable or repeatable. Pretty much. Or put another way: an actual stop watch is a very meaningful performance measurement tool.
- Joker_vD 2y agoMy point is, it's impossible to test everything in a reasonable timeframe. It would be much, much more convenient to know (call it "having an accurate theory") beforehand which approach will be faster. Imagine having to design electronics the way we design performant programs. Will this opamp survive the load? Who knows, let's build and try these five alternatives of the circuit and see which one of them will not blow. Oh, this one survived but it distorts the input signal horribly ("yeah, this one is fast, but it has multithreading correctness issues and reintroduction of locks makes it again about as slow"), what a shame. Back to the drawing board.