4 ms·
I find this claim hard to believe honestly,could you point to examples where performance is limited by Dram speed and not by cpu / caches? They must be applicat
by mda 5y ago
I find this claim hard to believe honestly,could you point to examples where performance is limited by Dram speed and not by cpu / caches? They must be applications with extremely bad design causing super low cache hits.
- staticassertion 5y ago> They must be applications with extremely bad design causing super low cache hits. So basically any program written in a language with pointer types exclusively.
- throwaway894345 5y agoI thought I’d heard that Java VMs go to great lengths to maintain cache coherence? I’d be curious to hear from the Lisp folks because I always hear that Lisps can be surprisingly performant.
- staticassertion 5y agoThere are some optimizations in the JVM that will improve cache coherence. Bump allocation helps, inlining and escape analysis help, etc. In theory the GC can also rearrange memory to 'compact' it. I'm not aware of this optimization in practice.
- throwaway894345 5y agoAs far as I know, all mainstream Java GCs are compacting; however, I don’t have an idea about the degree this improves cache coherence.
- mda 5y agoEven then, today's cpus have enormous caches, and not all parts of program is pointers. you cant make a crappy application much faster just because you have faster ram.
- erulabs 5y ago> They must be applications with extremely bad design causing super low cache hits Yep, this is exactly the case - also, on systems that are busy and context-switching often and thus flushing their cpu caches more frequently. Combine the two, busy systems running loads of un-optimized code, and boom, you have described how most computers run in the real world. This is why "synthetic" benchmarks, which are well designed code running on quiet machines more or less match up to CPU Frequency exclusively. I don't really have any good charts to show you, but you might checkout an old review of the processor I mentioned as having one of the first on-die memory controllers: https://techreport.com/review/5655/amds-opteron-146-processor/ https://techreport.com/review/5655/amds-opteron-146-processo... The AMD Opteron 240 1.4GHz keeps up with chips close to 2x it's frequency - and the memory access times are close to 1/2 as costly (ie: almost all the performance gain from 2x frequency is made up by 1/2 memory access time) - this makes sense, but remember these are well optimized applications (POV-Ray and Lightwave were extremely synthetic). In the real world, opening 10 misc windows applications from 2003, the K8 (particularly when overclocked) was a _beast_.
- sroussey 5y agoThis is why you dedicate entire machines to the same kind of load. All application code or all database. There was a moment where people tried to integrate—which was indeed faster but only for very limited use cases.
- mda 5y agoWell, Opteron is an ancient processor, I don't think we can make any conclusion based on that. Today's server processors has enormous caches compared to Opteron. Honestly in cases you mention, badly designed processes killing the Cpu, I fail to see how faster ram makes a huge difference.
- erulabs 5y agoI mean - the people who designed Gravaton 3 seem to agree with my premise, so at least that’s some validation. Alternatively, do some CPU profiling with your workstation - a massive amount of time is simple waiting for memory returns.
- therealcamino 5y agoAnything where the working set is larger than L3 cache.
- hexxagone 5y agoIn data compression, inverting a BWT with large blocks or using Context Mixing to compress large blocks (which requires huge context maps). These 2 cases require a lot of random memory accesses.