Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
dendibakh
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
dendibakh
3y ago
... And the only (?) memory profiler for Windows.
2.
▲
Performance Benefits of Using Pages for Code
(easyperf.net)
1 points
by
dendibakh
4y ago
|
0 comments
3.
▲
by
dendibakh
7y ago
Thanks, and don't forget to measure. :)
4.
▲
by
dendibakh
7y ago
You can also visit easyperf.net (my blog). If you like Linux perf the same way I do, you will like it. Here are some links: - https://easyperf.net/blog/2018/08/26/Basics-of-profiling-wit... - https:
5.
▲
A Programmer’s Guide to Performance Analysis and Tuning on Modern CPUs
(linkedin.com)
2 points
by
dendibakh
7y ago
|
0 comments
6.
▲
Everything you wanted to know about Intel Processor Traces
(easyperf.net)
1 points
by
dendibakh
7y ago
|
0 comments
7.
▲
by
dendibakh
7y ago
Thanks for the great comment. Where I can read more about the impact of each item on variability?
8.
▲
by
dendibakh
7y ago
Thanks for the comment. As answered to the previous comment, yes, I agree. Likely I did a poor job of explicitly saying that those advises are not generally applicable. It usually makes sense to have a dedicated machine (pool of machines) t
9.
▲
by
dendibakh
7y ago
Yep. And I think I stated that in the beginning of the article. :) This one: > It is important that you understand one thing before we start. If you use all the advices in this article it is not how your application will run in practice.
10.
▲
by
dendibakh
7y ago
Yes, that might be possible. However, you probably will get multiple cycle counts for the same function depending on which path was taken. And it works only if the amount of taken branches in the function is not that big (less than 32). Oth
11.
▲
by
dendibakh
7y ago
Thanks. I'm glad you like the article. :)
12.
▲
Precise timing of machine code with Linux perf
(dendibakh.github.io)
2 points
by
dendibakh
8y ago
|
0 comments
13.
▲
How to collect CPU performance counters on Windows?
(dendibakh.github.io)
1 points
by
dendibakh
8y ago
|
0 comments
14.
▲
Performance optimization contest
(dendibakh.github.io)
2 points
by
dendibakh
8y ago
|
0 comments
15.
▲
by
dendibakh
8y ago
Hi, I'm glad you like the article. The process how I went from TMAM metric to particular event that was used to calculate it is describe in the TMAM metrics table: https://download.01.org/perfmon/TMA_Metrics.xlsx
16.
▲
How good of a performance optimizer you are? Contest
(dendibakh.github.io)
1 points
by
dendibakh
8y ago
|
0 comments
17.
▲
Improving performance by better code locality
(dendibakh.github.io)
3 points
by
dendibakh
8y ago
|
0 comments
18.
▲
Advanced profiling topics. PEBS and LBR
(dendibakh.github.io)
2 points
by
dendibakh
8y ago
|
0 comments
19.
▲
PMU counters and profiling basics
(dendibakh.github.io)
1 points
by
dendibakh
8y ago
|
0 comments
20.
▲
by
dendibakh
9y ago
Well, yeah. I don't know much about the current state of the art (because I don't touch the CodeGen on a daily basis), but I kind of look into the future with hope that compilers will better handle at least those "simple"
21.
▲
by
dendibakh
9y ago
Thanks for this clear explanation! Regarding your example with 1000 muls followed by 1000 loads... That's why in my experiments I interleaved loads and bswaps, because that's what (hopefully) every decent compiler will do.
22.
▲
MicroFusion in Intel CPUs
(dendibakh.github.io)
1 points
by
dendibakh
9y ago
|
0 comments
23.
▲
by
dendibakh
9y ago
Hi, Thanks to everybody who commented the post. With your help I was able to close the gap in my knowledge and summarized it in my new post: https://dendibakh.github.io/blog/2018/02/15/MicroFusion-in-I...