4 ms·
Performance Hints
- jesse__ 9mo agoWonderful article. I wish more people had this pragmatic approach when thinking about performance
- menaerus 9mo agoI actually wish the audience to take the opposite, or perhaps a more balanced view. Being pragmatic is like taking an extreme view and as much as this article is a great resource, and contains some legit advice otherwise difficult to find elsewhere in such a concise form, folks need to be aware that this advice is what Google found for their unfathomable scale codebase to gain some real world benefits. The things this article is describing are more nuanced than just "think about the performance sooner than latter". I say this as someone who does these kind of optimizations for a living and all too often I see teams wasting time trying to micro-optimize codepaths which by the end of the day do not provide any real demonstrable value. And this is a real trap you can get into really easily if you read this article as a general wisdom, which is not.
- xnx 9mo agoThis formatting is more intuitive to me. L1 cache reference 2,000,000,000 ops/sec L2 cache reference 333,333,333 ops/sec Branch mispredict 200,000,000 ops/sec Mutex lock/unlock (uncontended) 66,666,667 ops/sec Main memory reference 20,000,000 ops/sec Compress 1K bytes with Snappy 1,000,000 ops/sec Read 4KB from SSD 50,000 ops/sec Round trip within same datacenter 20,000 ops/sec Read 1MB sequentially from memory 15,625 ops/sec Read 1MB over 100 Gbps network 10,000 ops/sec Read 1MB from SSD 1,000 ops/sec Disk seek 200 ops/sec Read 1MB sequentially from disk 100 ops/sec Send packet CA->Netherlands->CA 7 ops/sec
- barfoure 9mo agoThe reason why that formatting is not used is because it’s not useful nor true. The table in the article is far more relevant to the person optimizing things. How many of those I can hypothetically execute per second is a data point for the marketing team. Everyone else is beholden to real world data sets and data reads and fetches that are widely distributed in terms of timing.
- deleted 9mo ago[deleted]
- twotwotwo 9mo agoYour version only describes what happens if you do the operations serially, though. For example, a consumer SSD can do a million (or more) operations in a second not 50K, and you can send a lot more than 7 total packets between CA and the Netherlands in a second, but to do either of those you need to take advantage of parallelism. If the reciprocal numbers are more intuitive for you you can still say an L1 cache reference takes 1/2,000,000,000 sec. It's "ops/sec" that makes it look like it's a throughput. An interesting thing about the latency numbers is they mostly don't vary with scale, whereas something like the total throughput with your SSD or the Internet depends on the size of your storage or network setups, respectively. And aggregate CPU throughput varies with core count, for example. I do think it's still interesting to think about throughputs (and other things like capacities) of a "reference deployment": that can affect architectural things like "can I do this in RAM?", "can I do this on one box?", "what optimizations do I need to fix potential bottlenecks in XYZ?", "is resource X or Y scarcer?" and so on. That was kind of done in "The Datacenter as a Computer" (https://pages.cs.wisc.edu/~shivaram/cs744-readings/dc-computer-v3.pdf https://pages.cs.wisc.edu/~shivaram/cs744-readings/dc-comput... and https://books.google.com/books?id=Td51DwAAQBAJ&pg=PA72#v=onepage&q&f=false https://books.google.com/books?id=Td51DwAAQBAJ&pg=PA72#v=one... ) with a machine, rack, and cluster as the units. That diagram is about the storage hierarchy and doesn't mention compute, and a lot has improved since 2018, but an expanded table like that is still seems like an interesting tool for engineering a system.
- 9mo ago
- barfoure 9mo agoSome of this can be reduced to a trivial form, which is to say practiced in reality on a reasonable scale, by getting your hands on a microcontroller. Not RTOS or Linux or any of that, but just a microcontroller without an OS, and learning it and learning its internal fetching architecture and getting comfortable with timings, and seeing how the latency numbers go up when you introduce external memory such as SD Cards and the like. Knowing to read the assembly printout and see how the instruction cycles add up in the pipeline is also good, because at least you know what is happening. It will then make it much easier to apply the same careful mentality to this which is ultimately what this whole optimization game is about - optimizing where time is spent with what data. Otherwise, someone telling you so-and-so takes nanoseconds or microseconds will be alien to you because you wouldn’t normally be exposed to an environment where you regularly count in clock cycles. So consider this a learning opportunity.
- simonask 9mo agoJust be careful not to blindly apply the same techniques to a mobile or desktop class CPU or above. A lot of code can be pessimized by golfing instruction counts, hurting instruction-level parallelism and microcode optimizations by introducing false data dependencies. Compilers outperform humans here almost all the time.
- barfoure 9mo agoIt is not about outperforming the compiler - it’s about being comfortable with measuring where your clock cycles are spent, and for that you first need to be comfortable with clock cycle scale of timing. You’re not expected to rewrite the program in assembly. But you should have a general idea given an instruction what its execution entails, and where the data is actually coming from. A read from different busses means different timings. Compilers make mistakes too and they can output very erroneous code. But that’s a different topic.
- jesse__ 9mo agoExcellent corrective summary. "Compilers can do all these great transformations, but they can also be incredibly dumb" -Mike Acton, CPPCON 2014
- deleted 9mo ago[deleted]
- squirrellous 9mo agoI wish Google would open source their gtl library. Similar utilities exist elsewhere but not in the same consistent quality and well-integrated package. I particularly like the “what to do for flat profiles” ad “protobuf tips” sections. Similar advice distilled to this level is difficult to find elsewhere.
- canyp 9mo agoInteresting that the blog only runs until 2023. Have they been absorbed by AI, Rust, or both? I think I'd rather be eaten by a giant crustacean than work on AI.
- svat 9mo agoThe HN title here is currently “Performance Hints (2023)”, but this was only published externally recently (2025). (See e.g. https://x.com/JeffDean/status/2002089534188892256 https://x.com/JeffDean/status/2002089534188892256 announcing it.) And of course 2023 is when the document was first created, but much of the content is more recent than that. So IMO it's a bit misleading to put "(2023)" in the title.
- danlark1 9mo agoSurprisingly I didn't put 2023, it was merged with another submission possibly with the help of mods
- dredmorbius 9mo agoFYI: It's possible for that to be edited by others.
- zahlman 9mo agoIf the numbers come from analyzing performance in 2023, that seems more important than the external publication time.
- svat 9mo agoThe page is about tips for writing fast code. Much of it applied 20 years ago, and will apply 20 years from now. If by "the numbers" you mean specifically just the table ("rough costs for some basic low-level operations") in the "Estimation" section (which accounts for less than 0.5% of the words on the page), then that table was initially created in 2007, and is up-to-date as of 2025. Other numbers on the page are given with their dates, like 2001 and so on. So 2023 does not seem relevant in any way.
- dang 9mo ago(Ok, we've belatedly taken 2023 out of the title now)
- justicehunter 9mo agoReally helps to have all this good info in one page. I often find myself focusing on few aspects here while ignoring the rest. Definitely saved to remind myself that there is a lot more to performance that the few tricks I know.