11 ms·
Latency numbers every programmer should know
- dockd 14y agoDoes anyone feel like this is sort of an apples to oranges table? It compares reading one item from L1 cache to reading 1M byte from memory, without adjusting for the amount of data being read (10^6 more). It looks like the data was chosen to minimize the number of digits in the right column.
- Morg 14y agoSomeone should add basic numbers like ns count for 63 cycles modulo and that type of stuff - That'll help bad devs realize why putting another useless cmp inside a loop is dumb, and why alt rows in a table should NEVER be implemented by use of a modulo, for example. Yes I know that's not latency per se but in the end it is too.
- teach 14y agoI think if you're worried about whether or not you use modulo to calculate alternating table rows (and you don't work for Facebook), then you're almost certainly optimizing prematurely.
- Morg 14y agoIT DOES NOT COST MORE TIME TO CODE CORRECTLY Some approaches are NOT acceptable, it's not about optimizing prematurely, it's about coding obvious crap. While you may be used to the usual "code crap, fix later" and "waste cycles, there are too many of it" , it doesn't mean you're right. Everyone says it but you're still running on C (linux, unix), you're still going nuts over scaling issues (lol nosql for everyone) and you're still paying your amazon cloud bill.
- njs12345 14y agoOr, you know, just get a decent compiler: http://publications.csail.mit.edu/lcs/pubs/pdf/MIT-LCS-TM-600.pdf http://publications.csail.mit.edu/lcs/pubs/pdf/MIT-LCS-TM-60...
- Morg 14y agoI suppose you are referring to the very particular case of the right shift, but as much as that's easily predictable, it's a corner case. Who knows maybe the trend will be 3 colors instead of two. Or maybe it'll be another instruction that's wrongly abused. Or another compiler that actually sucks, like most JS interpreters. The idea really is to use the simplest logical approach to the problem rather than the wrong one. In the very well known case of the alt row table, it looks to me like we're alternating odd and even, why not just code that to start with, before any optimization ?
- njs12345 14y agoNo, this form of strength reduction can often eliminate modulo operations in a loop even when the modulus is not a constant. The example given is: for(t = 0; t < T; t++) for(i = 0; i < NN; i++) A[i%N] = 0; which is optimised to this, without a modulo in sight: _invt = (NN-1)/N; for(t = 0; t <= T-1; t++) { for(_Mdi = 0; _Mdi <= _invt; _Mdi++) { _peeli = 0; for(i = N*_Mdi; i <= min(N*_Mdi+N-1,NN-1); i++) { A[_peeli] = 0; _peeli = _peeli + 1; } } } I find the modulo easier to read in this case, but I guess that's a question of taste. It's certainly not 'wrong' to use a modulo, and probably worth the trade off in most cases if it makes your code clearer.
- Morg 14y agoYes, sometimes the compiler can compensate bad decisions from the programmer, the jvm can collect your garbage etc. - none of these will save you from stupid data models and idiotic objects.
- teach 14y agoI know you're ranting to the world at large, but I am not "going nuts over scaling issues". All my websites are static HTML files. I regenerate them as needed using custom Python code and my "databases", which are text files in JSON. I have several sites running on a single smallest Linode, and the CPU utilization virtually never cracks 1%. Also, note that I am not advocating "coding crap". I'm talking about not berating coworkers over the nanosecond cost of an extra modulo inside a loop.
- Morg 14y agoIf said coworkers are actually trying to improve and can take the advice peacefully, I will deliver it peacefully. The others I will be pleased not to work with.
- mseebach 14y agoIt does cost considerable time and brain bandwidth to learn to "code correctly" if coding correctly means knowing how to avoid every excess few nanoseconds. If your code is expressive, easy to reason about and fast enough, then less expressive, harder to reason about and even faster code isn't more correct.
- hythloday 14y agoQuite to the contrary, the optimization of using any particular method to colour rows is so tiny it can easily be outweighed over its lifetime by the 50 or so extra keystrokes it needs to type. That's how trivial this is (which is why people are reacting to your extremely aggressive tone).
- Morg 14y agoIndeed, I should drop the agression. However, the subject is not optimization but coding correctly in the first place. And the anti-optimization argument would be correct if: -typing represented more than 1% of dev work -code was never reused -code was never massively used -code had a short lifespan So let me help you see clearly: -I'm not a typist -Every bad code tutorial out there creates millions of code bits that contain the N times slower version, with an aggregate impact that actually matters -Any 10% opt mistake in a codebase like iptables would cause more carbon than you can imagine -Fortran is still in use because it's the fastest language there is with the best math libraries. Those seem to be eternal so far, and C seems to remain the only other relevant language throughout the short history of coding. Sure, there are much more problematic cases than the dumb even odd example, but I picked that one because many would recognize it.
- jeltz 14y agoActually it does not matter if you are facebook or not. What really matters is how tight the loop is and how much time is spent in it. EDIT: I agree with Morg. If coding right also results in faster code there is no reason not to do that.
- recursive 14y agoHow should they be implemented? And per se should NEVER be spelled "per say".
- Morg 14y agoIndeed it should never be spelled wrong, as it means in itself in latin, my bad really. Alt rows are a simple concept, the first row is odd, the next is even, etc. A good step forward is an if/then/else or a switch or an unrolled loop - a huge step forward in terms of performance too, as a mod takes 63 cycles and a cmp takes almost nothing. an example could be rowClass='even'; loop if(rowClass=='odd'){ rowClass='even'; }else{ rowClass='odd'; } endloop
- hythloday 14y agoI think I must be misreading you. Are you suggesting doing a string comparison to avoid the performance hit of a mod?
- Morg 14y agoI did write it like that yes. And it would still be faster than a mod, too, even though one byte might be better for registry usage, it won't affect cycles that much iirc.
- recursive 14y agoHey, guess what? You're wrong. (at least in python, which is a reasonable guess for a language that's generating html) >>> import timeit >>> timeit.Timer(stmt="z=101%2").timeit() 0.033080740708665485 >>> timeit.Timer(stmt="z='even'=='odd'").timeit() 0.05949918215862482
- hythloday 14y agoIt does seem to be true for javascript though: > profile = function(fn) { var start = Date.now(); fn(); return Date.now() - start; } > cmp = function() { for (var i=0; i < 1000000000; i++) { var z = 'odd' === 'even'; } } > mod = function() { for (var i=0; i < 1000000000; i++) { var z = 101 % 2; } } > prof(cmp) 20329 > prof(mod) 40792 Whether you think those 20 nanoseconds per test are worth saving is, I guess, an open question. :) I can imagine it being useful for game programming, for example.
- debacle 14y agoIn general, there are very few things you actually need a modulo for. It's highly inefficient.
- bitwize 14y agoA lot of compilers are smart enough these days to optimize modulo by n, n a power of 2, to bitwise AND the complement of n-1.
- yuvadam 14y agoIs a single-text-file-github-gist the best way to disseminate this piece of knowledge (originally by Peter Norvig, BTW)? What about a comprehensive explanation as to why those numbers actually matter? Meh.
- willvarfar 14y agohttp://www.infoq.com/presentations/Lock-free-Algorithms http://www.infoq.com/presentations/Lock-free-Algorithms early-on gives good numbers on-machine. I like this visualisation too: http://news.ycombinator.com/item?id=702713 http://news.ycombinator.com/item?id=702713
- matthavener 14y agoFor anyone thinking about watching: the presentation is incredibly good. The name is kinda misleading -- they talk more about modern x86 architecture and actual numbers of various algorithms than the lock free algorithms themselves. Both of those guys have a high emphasis on measuring and testing to improve performance.
- jgrahamc 14y agoI believe that this originally comes from Norvig's "Teach Yourself to Program in Ten Years" article: http://norvig.com/21-days.html http://norvig.com/21-days.html
- alecco 14y agoThat's from 2011. Amazon's James Hamilton wrote about it in 2009, and it comes from a presentation that year by Google's Jeff Dean: http://perspectives.mvdirona.com/2009/10/17/JeffDeanDesignLessonsAndAdviceFromBuildingLargeScaleDistributedSystems.aspx http://perspectives.mvdirona.com/2009/10/17/JeffDeanDesignLe... EDIT: this is wrong, Norvig's page pre-dates Dean's presentation http://wayback.archive.org/web/*/http://norvig.com/21-days.html http://wayback.archive.org/web/*/http://norvig.com/21-days.h...
- tjr 14y agoI make no claim as to who first assembled that particular table of data, but Norvig's article is dated 2001.
- alecco 14y agoEDIT: I stand corrected. The numbers seem to have been evolving and the original source seems to be that page. http://wayback.archive.org/web/*/http://norvig.com/21-days.html http://wayback.archive.org/web/*/http://norvig.com/21-days.h...
- luckydude 14y agoThis is shameless self promotion (well, me and Carl promotion) but we were measuring these sorts of things in the late 80's and wrote a paper about it that got best paper at Usenix in 95. I think we had most of those numbers, not in the same format. Pissed off the BSD folks because it made them look bad. Oh, well. Helped make Linux better, largely because while the BSD guys refused to engage, Linus did. He and I spent many many hours discussing what was the right thing to measure and what should not be measured. We both felt that lmbench would influence OS design (and it's influenced processor design, see all the cache prefetch stuff, I'm pretty convinced that's because all the processor people used lmbench). Linus was already on the "OS should be cheap path" but lmbench helped him make the case to other people who wanted to add overhead because of their pet project. The cool part about working with Linus was he was never about making Linux look better, he was about measuring the right things. If Linux sucked, oh, well, he'd fix it or get someone else to fix it. Awesome attitude, I feel the same way. The only published work that might predate lmbench for these sorts of numbers is Hennessy and Patterson computer architecture. They talked about memory latency but so far as I recall, didn't have a benchmark. That said, that book is friggin awesome and anyone who cares about this sort of thing and hasn't carefully read that book is missing out.
- Symmetry 14y agoIn practice any out of order processor worth its salt ought to be able to entirely hide L1 cache latencies.
- matthavener 14y agoAgreed, I think the real lesson is: reading is 10x cheaper than branching, so if you can do something with 10 non-branching ops it'll be just as fast as a single branch.
- seabee 14y agoConversely: if you can do something with a branch that's correctly predicted 90% of the time it'll be just as fast as 10 non-branching ops. Branch prediction is a tool like any other - don't neglect it when it can help you.
- marshray 14y agoAs a coder, how can I predict the branch predictor?
- CamperBob2 14y agoThere are definite rules, documented by the CPU vendor. A forward branch is assumed not taken, while a backward branch is assumed to come at the end of a loop that will probably iterate more than once. See http://software.intel.com/en-us/articles/branch-and-loop-reorganization-to-prevent-mispredicts/ http://software.intel.com/en-us/articles/branch-and-loop-reo... for example. I'd assume the particulars will vary between CPU manufacturers and families, but the idea that backward branches will probably be taken seems fairly universal.
- haberman 14y agoThis is outdated information; Intel chips have not used static prediction for conditional branches since NetBurst (Pentium 4). "Pentium M, Intel Core Solo and Intel Core Duo processors do not statically predict conditional branches according to the jump direction. All conditional branches are dynamically predicted, even at first appearance." --http://www.intel.com/content/dam/doc/manual/64-ia-32-architectures-optimization-manual.pdf http://www.intel.com/content/dam/doc/manual/64-ia-32-archite...
- larsberg 14y agoHonestly, I'd rather programmers know how to _measure_ these numbers than just have them memorized. I mean, if I told them that their machine had L3 cache now, what would they do find out how that changes things? (This comment is also a shameless plug for the fantastic CS:APP book out of CMU).
- nhebb 14y ago+1 to CS: APP (Computer Systems: A Programmer's Perspective, by Bryant and O'Hallaron). I thought Computer Architecture: A Quantitative Approach by Hennessey was dry and better suited to EE's. CS: APP, on the other hand, really sucked me in - and I'm not the kind of guy that typically gets enthralled by CS texts.
- apaprocki 14y agoDTrace lets you measure L# cache hits/misses along with lots of other useful things.
- adobriyan 14y agoIt lets you measure a number of cache hits, not the latency.
- memset 14y agoHonest question: how would you measure an L2 cache lookup? (What program would I need to write which ensures that a value is stored in L2, such that when I read it later, I would know the lookup time to indeed be the L2 time?) At that, how would one measure this kind of thing at the nanoseconds level? Would using C's clock_gettime() functions be good enough? Is there any other facility to count cycles which have passed between two operations?
- Swizec 14y agocracks knuckles Let's see if I remember enough of what I'm supposed to know in an exam in a couple of months or so. You need to know the size of a page in your L1 cache. Then you can write a program that goes through a big enough chunk of memory that an "out of space" error occurs predictably in L1 cache, but not in L2 cache. That way you can know when specifically your program went to L2 cache to swap out a page of L1 cache when it was needed. The caveat being that you can probably only do this well enough in some sort of assembler code (for your architecture) and that you would have to be running a single-process system without interrupts enabled. Otherwise all sorts of things can mess up your cache lookups.
- peteretep 14y agoTook me forever to find this, but: https://plus.google.com/112493031290529814667/posts/LvhVwngPqSC https://plus.google.com/112493031290529814667/posts/LvhVwngP...
- hellerbarde 14y agoFrom that link, this brilliant visualisation: http://i.imgur.com/X1Hi1.gif http://i.imgur.com/X1Hi1.gif
- al_james 14y agoWhat stands out here is how long a disk read takes (especially compared to network latency). Indeed, disk is the new tape.
- ryandetzel 14y agoInteresting but unnecessary for most programmers today. I'd rather my programmers know the latency of redis vs memcached vs mysql and data type ranges.
- snotrockets 14y agoEvery programmer that understands (rather than "knows") the number mentioned in the link would know the numbers you are looking for (or at least, the relationships between those.) I'm not sure this relation holds for the opposite direction.
- aristus 14y agoThese are good rules of thumb, but need more context. Plugging an article I wrote about this & other things a couple of years ago for FB engineering: https://www.facebook.com/note.php?note_id=461505383919 https://www.facebook.com/note.php?note_id=461505383919 The "DELETE FROM some_table" example is bogus, but the rest is still valid.
- kjhughes 14y agoAnyone who hasn't heard Rear Admiral Grace Murray Hopper describe a nanosecond should check out her classic explanation: http://www.youtube.com/watch?v=JEpsKnWZrJ8 http://www.youtube.com/watch?v=JEpsKnWZrJ8
- DanBC 14y agoThe long side of a piece of A4 paper is 297 mm. Light can travel, in one nano second, 299.8 mm.
- stiff 14y agoI present to you Grace Hopper handing people nanoseconds out: http://www.youtube.com/watch?v=JEpsKnWZrJ8 http://www.youtube.com/watch?v=JEpsKnWZrJ8 :)
- CookWithMe 14y agoWhat about L3 Cache? What about Memory Access on another NUMA Node? What about SSD? Does a mobile phone programmer need to know the access time for disks? Does an embedded system programmer need to know anything of these numbers? Every programmer should know what memory hierarchy and network latency is. (If you learn it by looking at these numbers, fine...)
- jbooth 14y agoI'm not an expert, but: L3 is generally on the order of the same time as main memory - it's main purpose is to reduce the total amount of requests in order to conserve bandwidth SSDs are on the order of 0.1ms, so 100,000 ns, give or take a factor of 10. Someone smarter than me will have to answer the NUMA node question.
- MichaelGG 14y agoI think L3 is much faster; perhaps 1/3rd of main memory, assuming the line is available and not in another core. Here are some numbers for L3 cache, from Intel (probably specific to the 5500 series)[1]: L3 CACHE hit, line unshared ~40 cycles L3 CACHE hit, shared line in another core ~65 cycles L3 CACHE hit, modified in another core ~75 cycles remote L3 CACHE ~100-300 cycles Local Dram ~60 ns Remote Dram ~100 ns 60ns at 2.4GHz is ~144 cycles, right? 1: http://software.intel.com/sites/products/collateral/hpc/vtune/performance_analysis_guide.pdf http://software.intel.com/sites/products/collateral/hpc/vtun...
- CookWithMe 14y agoSorry, these were rhetorical questions :) I was trying to make the point that these numbers are somewhat arbitrary (i.e. why do I need to know the access speed to disc, when I keep everything in memory on a NUMA system?) and don't apply to all programmers (e.g. embedded systems may not have discs, L2 Caches or internet access).
- zippie 14y agoThese numbers by Jeff Dean are relatively true but need to be refreshed for modern DRAM modules & controllers. Specifically, the main memory latency numbers are more applicable to DDR2 RAM vs the now widely deployed DDR3/DDR4 RAM (more channels = more latency). This has been a industry trend for a while and theres no change on the horizon. Additionally, memory access becomes more expensive because of CPU cross chatter when validating data loads across caches. A potential pitfall with these numbers is they give engineers a false sense of security. They serve as a great conceptual aid - network/disk I/O are expensive and memory access is relatively cheap but engineers take that to an extreme, and get lackadaisical about memory access. When utilizing a massive index (btree) our search engine failed to meet SLA because of memory access patterns. Our engineers tried things at the system (numa policy) and application level (different userspace memory managers, etc.) Ultimately, it all came down to improving the efficiency around memory access. We used Low-Level Data Structure to get the 2x improvements in memory latency: https://github.com/johnj/llds https://github.com/johnj/llds
- lallysingh 14y agoIf this is your cup of tea, have a look at Agner Fog's resources: http://agner.org/optimize/ http://agner.org/optimize/ Also, I'd have a look at Intel's VTune or the 'perf' tool that ships with the linux kernel.
- dsr_ 14y agoScaling up to human timeframes, one billion to one: Pull the trigger on a drill in your hand 0.5s Pick up a drill from where you put it down 5s Find the right bit in the case 7s Change bits 25s Go get the toolkit from the truck 100s Go to the store, buy a new tool 3000s Work from noon until 5:30 20000s Part won't be in for three days 250000s Part won't be in until next week 500000s Almost four months 10000000s 8 months 20000000s Five years. 150000000s
- hellerbarde 14y agoI forked the gist and did something similar. Would you mind if I took some of your suggestions into my fork? https://gist.github.com/2843375 https://gist.github.com/2843375
- dsr_ 14y agoFine by me.
- CamperBob2 14y agoThis is the same basic advice that I try to keep in mind when writing English text. A word that the reader already knows is in "L1." A word they have to stop and think about is an L1 miss. If they have to reach for the dictionary, it's an L2 miss -- and probably several more cycles for a line fill, as they get distracted reading the next few entries in the dictionary. As programmers, Strunk & White might have made the list of the all-time greats.
- EternalFury 14y agoConsidering that so many programmers are currently enthralled with JavaScript, Ruby, Python and other very very high level languages, the top half of this chart must look very mysterious and unattainable.
- luckydude 14y agoMost of these latencies were measured and written up for a bunch of systems by Carl Staelin and I back in the 1990's. There is a usenix paper that describes how it was done and the benchmarks are open source, you can apt-get them. http://www.bitmover.com/lmbench/lmbench-usenix.pdf http://www.bitmover.com/lmbench/lmbench-usenix.pdf If you look at the memory latency results carefully, you can easily read off L1, L2, L3, main memory, memory + TLB miss latencies. If you look at them harder, you can read off cache sizes and associativity, cache line sizes, and page size. Here is a 3-D graph that Wayne Scott did at Intel from a tweaked version of the memory latency test. http://www.bitmover.com/lmbench/mem_lat3.pdf http://www.bitmover.com/lmbench/mem_lat3.pdf His standard interview question is to show the candidate that graph and say "tell me everything you can about this processor and memory system". It's usually a 2 hour conversation if the candidate is good.
- ajross 14y agoPedantic quip: I have a hard time believing you guys were measuring half nanosecond cache latencies on a machine with a 100MHz clock. :) And actually the cache numbers seem optimistic, if anything. My memory is that a L1 cache hit on SNB is 5 cycles, which is 2-3x as long as that table shows.
- luckydude 14y agoWe didn't believe it either until we put a logic analyzer on the bus and found that the numbers were spot with respect to the number of cycles. I don't remember how far off they were but it wasn't much, all the hardware dudes were amazed that software could get that close. tl;dr: the numbers were accurate to the # of cycles, might have been as much as 1/2 of 1 cycle off. Edit: I should add this was almost 20 years ago, I dunno how well it works today. Sec, lemme go test on a local machine. OK, I ran on a Intel(R) Core(TM) i7-3930K CPU @ 3.20GHz (I think that's a Sandy Bridge) that is overclocked to 4289 MHz according to mhz, and it looks to me like that machine takes 4 cycles to do a L1 load. That sound right? lmbench says 4.05 cycles. I poked a little more and I get L1 4 cycles, ~48K L2 12 cycles, ~256K L3 16 cycles, ~6M Off to google and see how far off I am. Whoops, work is calling, will check back later.
- 14y ago
- bunderbunder 14y agoI find myself thinking of figures like these every time I see results for benchmarks that barely touch the main memory brought up in debates about the relative merits of various programming languages.
- hobbyist 14y agoHow is mutex lock/unlock different from any other memory access?
- wmf 14y agoIt requires more cache coherence.
- SeanLuke 14y agoHow is a mutex lock less expensive than a memory access? Are such things done only in registers nowadays? This doesn't sound right.
- haberman 14y agoLess expensive than a main memory access. It must be that an uncontended lock/unlock can happen in cache.
- JoeAltmaier 14y agoNo definitely not. A full memory fence surrounds lock/unlock.
- scott_s 14y agoI actually don't think that's true. My understanding is that on x86, atomic instructions have implicit lock instructions before them. (Or you can make some instructions atomic by putting a lock instruction before them.) Such instructions lock the bus and prevent other cores or SMT threads from accessing memory. In that way, you can safely perform an atomic operation on a value in the cache. Note that this implies that atomic operations slow down others cores and SMT threads.
- JoeAltmaier 14y agoLocks are often implemented using an xchg instruction, which is implicitely locked. All processor's caches are committed/flushed for the affected cache line. So its correct to say other processors are slowed down. But it also in that sense IS a main memory operation, just not yours.
- scott_s 14y agoTo be clear, then we agree that haberman was correct, and the value can be changed in cache.
- CookWithMe 14y agoAlso, these numbers don't mean much on their own. E.g. L2 Cache is faster than main memory, but that doesn't help you if you don't know how big your L2 Cache is. Same for main memory vs. disc. E.g. I optimized a computer vision algorithm for using L2 and L3 caches properly (trying to reuse images or parts of images still in the caches). Started off with an Intel Xeon: 256KB L2 Cache, 12MB L3 Cache. Moved on to an AMD Opteron: 512KB L2 Cache (yay), 6MB L3 Cache (damn). Also, the concept of the L2 Cache has changed. Before multi-cores it was bigger and the last-level-cache. Now it has become smaller and the L3 Cache is the last-level-cache, but has some extra issues due to the sharing with other cores. The important concepts every programmer should know are memory hierarchy and network latency. The individual numbers can be looked up on a case-by-case basis.
- patrickmay 14y agoWhen working on low latency distributed systems I more than once had to remind a client that it's a minimum of 19 milliseconds from New York to London, no matter how fast our software might be.
- sciurus 14y agoOne of my favorite writeups on this topic is Gustavo Duarte's "What Your Computer Does While You Wait" http://duartes.org/gustavo/blog/post/what-your-computer-does-while-you-wait http://duartes.org/gustavo/blog/post/what-your-computer-does...
- JoeAltmaier 14y agoNetwork transmit time is almost irrelevant. It takes orders of magnitude more time to call the kernel, copy data, and reschedule after the operation completes than the wiretime. This paradox was the impetus behind Infiniband, virtual adapters, and a host of other paradigm changes that never caught on.
- Peaker 14y agoInfiniband has caught on, at least in the high-end.
- luckydude 14y agoHuh. Data please. Part of the reason I wrote lmbench was to make sure that what you are saying is not true. And it is not in Linux, kernel entry and exit is well under 50 nanoseconds. Passing a token back and forth, round trip, in an AF_UNIX socket is 30 usecs. A ping over gig ether is 120 usecs. Unless I'm completely misunderstanding, you are saying that the OS overhead should be "orders of magnitude" more than the network time, that's not at all what I'm seeing on Linux. I guess what you are saying is that given an infinitely fast network, the overhead of actually doing something with the data is going to be the dominating term. Yeah, true, but when do we care about the infinitely fast network in a vacuum? We always want that data to do something so we have to pay something to actually deliver it to a user process. Linux is hands down the best at doing so, it's not free but it is way closer to free than any other OS I've seen.
- Getahobby 14y agoThis may be ignorant and please correct if I am off base but given the same physical medium isn't sending 2k across the network the same cost whether you are at fastE or gigE? Given the network is not saturated?
- luckydude 14y agofastE is 100Mbits/sec, gigE is 1000Mbits/sec, so given the same size packet, gigE is in theory 10x faster. However, to make things work over copper I believe that gigE has a larger minimum packet size so it's not quite apples to apples on pings (latency). For bandwidth, the max size (w/o non-standard jumbo grams), is the same, around 1500 bytes, and gigE is pretty much linear, you can do 120MB/sec over gigE (and I have many times) but only 12MB/sec over fastE.
- mmukhin 14y ago2kB over 1Gbps is actually 16ns (i guess they round up to 20)
- some1else 14y agoJohn Carmack recently used a camera to measure that it takes longer to paint the screen in response to user input, than send a packet accross the Atlantic: http://superuser.com/questions/419070/transatlantic-ping-faster-than-sending-a-pixel-to-the-screen/419167#419167 http://superuser.com/questions/419070/transatlantic-ping-fas... I came across the post when I was looking for USB HID latency (8ms).
- perlpimp 14y agoSuch items are important to web developers and they can use them to justify looking at one or other technology. Or perhaps attempt at least benchmarking and have them as one of the guides in configuring and setting up services. Comes to mind why is Redis can be better then mongodb and in what configuration. As well in discussion about this and that these can be of help too. Adding misaligned memory penalties such as on word boundary and page boundary can enhance such document. This might be a good cheatsheet if one inclined to research and make one.
- xb95 14y agoThis reminds me of one of the pages that Google has internally that, very roughly, breaks down the cost of various things so you can calculate equivalencies. As an example of what I mean (i.e., these numbers and equivalencies are completely pulled out of thin-air and I am not asserting these in any way): * 1 Engineer-year = $100,000 * 25T of RAM = 1 Engineer-week * 1ms of display latency = 1 Engineer-year This allows engineeers to calculate tradeoffs when they're building things and to optimize their time for business impact. E.g.: it's not worth optimizing memory usage by itself, Latency is king, Don't waste your time shaving yaks, etc etc.
- Morg 14y agoYet everyone uses much slower RAM in servers and will likely continue to do so, all the while caches swell, etc. Optimizing memory usage is almost irrelevant today, until it starts being a bandwidth problem, and that's still solvable but only through complex scaling strategies that also cost several engineer-years.
- balloot 14y agoThe thing here that is eye opening to me, and relevant to any web programmer, is that accessing something from memory from another box in the same datacenter is about 25x times as fast as accessing something from disk locally. I would not have guessed that!
- hboon 14y agoThat's why tools like memcached work like they do.
- chmike 14y agoLz4 is faster than zippy and much easier to use. It's a single .h .c file.