4 ms·
One big difference is memory availability; the amount of RAM directly accessible from a 64-core CPU is much larger than that of the thousands of CUDA cores on a
by staticfloat 6y ago
One big difference is memory availability; the amount of RAM directly accessible from a 64-core CPU is much larger than that of the thousands of CUDA cores on a Quadro card. Perhaps for many workloads intelligent use of the PCIE bus can make up for that, streaming datasets in/out of the card, but for others, having random access to hundreds of GBs of data may be non-negotiable.
I'd be interested to know what workloads truly require that; most simulation solvers that I can think of that require large amounts of RAM are working on a discretized spatial grid, and the physical laws they're simulating are fairly "local"; e.g. you can process one area then move on to the next, in a nice stencil operation. This lends itself well to both parallelization and a sort of data-locality, so you can stream the necessary pieces of data in to the cores in an orderly fashion. I don't know what kinds of massively parallelizable workloads require much more random access of memory.
- choeger 6y agoDAE and ODE simulators do not work with that spatial distribution.
- vosper 6y agoI’m not sure about “require”, but we use multi-hundred-gigabytes or RAM instances for our Elasticsearch coordinator nodes, which seems to work well.
- CSDude 6y agoElasticsearch recommends <32GB because of 32-bit per their doc, is it different for coordinator nodes?
- adamtulinius 6y agoThat recommendation is only for the JVM heap to my knowledge.
- Xylakant 6y agoFor coordinator nodes, only the HEAP is relevant, excess RAM for disk buffers is only of interest for data holding nodes.
- Xylakant 6y agoThe underlying reason is that at around 32GB the JVM internally shifts from 32bit pointers for HEAP management to 64bit. The feature is called compressed object pointers (compressed oops) The increased pointer size uses substantial amounts of RAM, a reasonable estimate is that you need at least ~45GB to not have a loss in total available HEAP. Also, GC cycles get longer the more HEAP you have. However, some workloads may require more heap to even be possible and there’s pauseless garbage collector implementations able to handle substantial amounts of memory, so it is possible. It’s just not something you should be doing unless you know exactly what you are doing. See https://wiki.openjdk.java.net/display/HotSpot/CompressedOops https://wiki.openjdk.java.net/display/HotSpot/CompressedOops
- reitzensteinm 6y agoLacking value types, larger pointers will probably penalize JVM programs more than their peers.
- SomewhatLikely 6y agoWhile there are no JEP 169 style Value Types, you can still layout your data in a way that significantly reduces the number of pointers. For example using parallel arrays of primitives. It's not as clean, but it is one mitigation.
- behohippy 6y agoA "perfect" ES node has 64 gig of memory. 32 to heap and 32 for off heap as well as 8-12 cores for executing queries. Heap is mostly used for internal operations and the inverted index stuff (fielddata), off-heap was used for the columnar data store (docvalues). Most server class machines exceed this nowadays, so we would recommend containerizing multiple instances of ES on a single machine and using the rack awareness feature to make sure shards didn't replicate onto the same physical host. It's been a while since I worked at Elastic, so this info might be out of date by now, but we're still using this formula for our ECE setups where I work now.
- tromp 6y agoThe CuckatooN Proof-of-Work puzzle [1] requires 2^N bits of randomly accessed memory for efficient solving which parallelizes very well. In practice the random access is so slow that an alternative solver using 32 times more memory and mostly sequential access is much faster. [1] https://github.com/tromp/cuckoo https://github.com/tromp/cuckoo