12 ms·
7nm AMD EPYC “Rome” CPU with 64C/128T to Cost $8K (56 Core Intel Xeon: $25K-50K)
- deleted 7y ago[deleted]
- gigatexal 7y agoIf you need AVX/AVX2/AVX512 then go intel. Otherwise this AMD chip would be my choice for most applications.
- rat9988 7y agoI'm not sure intel has advantage in avx 2 with the new ryzen. That said I got this information from some unsourced comment in hacker news, it might be totally wrong.
- agumonkey 7y agodoes avx512 bring that much of a perf advantage ?
- gigatexal 7y agoFor software that takes advantage of it it does.
- continuations 7y agoAny commonly used software that benefits from AVX-512?
- bitL 7y agoVideo editing, audio mastering, games.
- reitzensteinm 7y agoGames!? No Man's Sky had trouble with SSE 4.1 being a requirement, and ended up patching it out... let alone AVX, AVX2, or AVX512. Typically, games won't take advantage of new instruction sets until it's ready to be a minimum requirement, as otherwise you need to maintain two code paths, the benefit of which is only to be speeding up execution on what are already the more powerful CPUs. What games can benefit from AVX512?
- bitL 7y agoNot now obviously as it is not mainstream yet, but AVX512 makes developer's life way easier than older AVX sets. So I see no reason why game engine developers wouldn't use them in the future. For now only some server-specific workloads use AVX512.
- reitzensteinm 7y agoIt's going to be a long, long time - at least a decade. I'd go as far as to say "never" is plausible, depending on how ARM penetration plays out. If the consoles go ARM before AVX-512 has been standard on new PCs for years, it may just not be worth it.
- throwaway2048 7y agoExcept it dosent because there are 4+ mutually incompatible versions of AVX instruction sets. Gotta love the price segmentation.
- gigatexal 7y agoAnd AVX512L should further improve that.
- mhh__ 7y agoSome LLVM compilers produce code that (possible with GCC too) can jit itself on a per function on first execution e.g. -march=native so no branches required. This requires good planning of the game, i.e. this reduces inlining so one needs to be careful with putting the important bits in the JITted code
- gigatexal 7y agoNot an authoritative list but a graphics benchmark and an x264 encoder at least according to this https://forums.anandtech.com/threads/avx-512-what-software-use-avx-512.2520858/ https://forums.anandtech.com/threads/avx-512-what-software-u...
- mda 7y agoRarely. It has serious throttling issues as well.
- agumonkey 7y agoI appreciate your tautological answer. I was more looking at concrete examples with % differences in benches :)
- namibj 7y agoWell, it includes some masking instructions that allow algorithms to use vectorization in cases where it's not natively processing vectors. The benefit should be similar to the thread-local control flow stack introduced in the Volta microarchitecture of Nvidia.
- deleted 7y ago[deleted]
- reitzensteinm 7y agoI'd say it's a bit early to be making statements like that; Zen 2 has native 256 bit vector units, and might yet do OK with AVX2 code vs Intel processors. AVX512 though, even if you don't use the full width registers (eg to avoid throttling), the new instructions are very useful and not at all present on Zen 2. I wouldn't be surprised if Zen 3 kept 256 bit vector units, but supported the AVX512 instruction set.
- bitL 7y agoMoreover, AVX2 operates at full frequency on Zen 2. AVX-512 heavily downclocks CPU; not sure about AVX2 loads on Intel right now (some people complained about AVX2 throttling on Skylake X).
- celrod 7y agoI thought I read somewhere that Zen 2 will downclock based on heat, rather than simple instruction set heuristics. I can't find a source though, so take that as hearsay for now. Intel CPUs do downclock for both avx and avx-512. Many motherboards let you configure that in the bios. I bought both a 9940X and a (delidded) 7980XE for avx-512 intensive workloads. I use a water cooler with a 360mm radiator. The bios for these (overclockable) chips contains "AVX Offset" and "AVX512 Offset" parameters. I haven't really tested avx loads, but the avx512 downclock is necesary. They run at 70-80C when running all cores at 3.5GHz in avx512 heavy loads. I don't want to push the temperatures further than that. I don't think there's any practical way to avoid having to downclock. It does gives a sizeable speed boost overall (for those workloads). Re: this thread I'd bet those avx512 workloads will be faster as avx2 workloads running on two 64 core CPUs with avx2 than one 56 core CPU with avx512, all else equal. But it sounds like, instead of all else being equal, things like IPC favor Zen2. EDIT: If folks happen to be interested, I compared a bunch of different "^" (aka, "pow") functions in Julia, running on my 9940X here: https://discourse.julialang.org/t/slow-arbitrary-base-exponentiation-a-b/25386/19 https://discourse.julialang.org/t/slow-arbitrary-base-expone... Someone else shared results with their Ryzen 2950X here: https://discourse.julialang.org/t/workstation-advice-for-mostly-julia-use/25395/13 https://discourse.julialang.org/t/workstation-advice-for-mos... The vectorized versions were those with "sleef" or "xsimd" in their name. They tended to be 1.5 to 3 times faster, while the nonvectorized versions were 1.25 to 1.35 times faster on the 9940X. Some of my other code is likely to show a much bigger difference. For example, many small matrix multiplication operations get to take advantage of avx512's masks to vectorize efficiently, as well as the fact avx512 has 32 instead of 16 floating point registers to hold larger matrix blocks in registers, increasing the vfma to vmov ratio. I suspect the 3.2x difference in speed in the "jsleefpowcob!" benchmark is because of the register counts. I suspect with avx512 the compiler was able to avoid register spills, while with avx2 it had to reload a lot of data on each loop iteration. The biggest problem with avx512 IMO is that compilers seem bad at taking advantage of it (eg, they never use masks) unless you babysit them / write code with vectorization constantly in mind. gcc for example will not use 512 bit vectors by default. You must explicitly specify "-mprefer-vector-width=512". My tests (mostly just the Polyhedron Fortran benchmarks, as a set of numerical code) seemed to confirm that gcc (gfortran) was doing the right thing. Meaning unless you intend to go low level and use it yourself (which can be a rewarding hobby!), or have workloads where optimized libraries exist, you won't see any benefit from avx512.
- tpetry 7y agoDidn‘t intel start segmenting their processor on avx support? So only the expensive processors get avx512?
- eisa01 7y agoWhat software is typically able to take advantage of these instructions? My use case is solving large LP/MIP problems for power markets, do those algorithms benefit?
- ulzeraj 7y agoNot serve-y but RPCS3 benefits a lot from those.
- namibj 7y agoDepends on the structure of your problem. Some lend themselves well to vectorization, others not nearly as much. Or at least the needed vectorized solvers are not yet available for some problem structures. Try to use perf-tools to determine how you're currently using these execution units, and then you can look how much it might help. Rule of thumb: double the bitwidth of vector instructions gets you 80% more speed, instead of the theoretical 100%.
- emanuensis 7y agoThose languages where the tensor/array is the fundamental unit: such as the APL family. J and K have optimized builds using AVX.
- mda 7y agoNo, bad idea. Test or check reviews for your common workloads first, then choose. New Zen cores quite powerful and there are many gotchas with AVX execution. In most cases I can see AMD would be better overall or offer good value.
- gigatexal 7y agoI agree and I plan to.
- rbanffy 7y agoIndeed. With any microarchitecture (or ISA) change (and Intel -> AMD is a huge architectural one) you need to test your workloads. You'll find all sort of odd and surprising performance differences caused by throttling, the way caches behave, the way libraries choose code paths, various inter-core latencies, and so on. At least this is a move that doesn't require a full recompile.
- themgt 7y agoBack in the early 2000s Apple was getting crushed because Moto/IBM's PowerPC couldn't keep up with the 880lb gorilla Intel. Who could have predicted even a few years ago Apple would wind up in a similar situation, yet this time for having bet on Intel? Could the 8 core Mac Pro have been $3650 with a $650 Rome CPU, vs. the reality of $6k with a $3k Xeon?
- leemailll 7y agoI suspect apple will release comps with amd cpus, but hackintosh with amd is already good to go
- ulzeraj 7y agoThey are stuck with intel because of thunderbolt.
- leemailll 7y agoactually Asrock release X570 for ryzen2 with PCIE4 and thunderbolt 3 (https://www.anandtech.com/show/14455/asrock-x570-aqua-heaviest-flagship-motherboard-ever-with-thunderbolt https://www.anandtech.com/show/14455/asrock-x570-aqua-heavie...).
- ulzeraj 7y agoNow that’s some great news. I want an AMD Apple system now.
- reitzensteinm 7y agoNote that the 1P variant costs only $4,955, or under $80 per core. The consumer chips cost roughly $50/core, making this the flattest ramp I can personally remember.
- yazr 7y agoDoes anyone have a graph of recent $/transistor progress? Its obvious that moore's law is stagnant, and clock speed is dead, etc, etc. But are we at least getting the same core size at a lower cost ?
- 0815test 7y agoI'm interested in this as well. Of course from a practical POV, especially with escalating NRE costs at recent nodes, we're clearly at a point where "getting the same core size at lower cost" is not going to be a thing unless you can afford to amortize that NRE over a huge volume of chips.
- ksec 7y ago>Does anyone have a graph of recent $/transistor progress? It wouldn't matter. You can't look at cost / transistor without factoring in Die Size and yield. Not to mention the cost of Higher Performance and Low Power Transistor are different. And the wafer price ( Cost / Transistor ) also exclude all the design cost and tooling around the node.
- ksec 7y agoThat is the brilliance of Chiplet strategy. Those 74nm 7nm Die are extremely cost effective. ( For reference they are even smaller than the SoC used in iPhone ) The 14nm I/O die ( 400mm2 ) seems to be most expensive part. I wonder what cost improvement could be done in that area. And Note: That 5K CPU has 256MB of L3 Cache. May be I could Run the OS not in RAM, but in Cache.
- 0815test 7y agoNote that Intel has also announced that they're going to use chiplets w/ 3D stacking for newer designs, starting in 2019. So this isn't something that's exclusive to AMD.
- rolleiflex 7y agoI do have an AMD Hackintosh, I don't see any reason you wouldn't be able to use these as a dev Hackintosh machine. It's gorgeously fast with my Ryzen 2700x, eats compilation time like no one's business and it's stable. And it's something around 1/4th the price of the same horsepower from native Intel Macs. I can't imagine the cost delta of a EPYC Rome Hackintosh to its Intel counterpart. It's not like I cheaped out either, I have the same cheese-grater like anodised dark grey aluminium case with impeccable internal structure for easy upgrades. (Link to case: https://www.youtube.com/watch?v=gcNsHS2U8RM https://www.youtube.com/watch?v=gcNsHS2U8RM)
- bitL 7y agoDo you have drivers for everything? Do you think it would work on a Threadripper with many NVidia GPUs as well?
- leemailll 7y agoMany nvidia card don't have driver for mojave or catalina. No 9xx nvidia gpu or later on can work
- pmjordan 7y agoCorrect, the built-in drivers only cover the "Kepler" architecture, as that was the last one that Apple used in a Mac. The 3rd party Nvidia drivers which support more modern GPUs are only offered up to macOS 10.13 High Sierra. The technical background for this is that Metal and OpenGL acceleration requires the GPU driver bundle to be loaded into a process's address space. (effectively a dynamic library) With 10.14, Apple has enabled the "library-validation" codesigning flag just about everywhere, which means that only libraries signed by the same developer as the process's main executable or by Apple can be loaded into a process. Hence, 3rd party GPU drivers won't load in WindowServer for example, which makes them dead in the water. Apple could of course implement an exception or sign Nvidia's driver, but so far it looks like they're not interested. Perhaps there will be a solution to coincide with release of the new Mac Pro which once again contains PCIe slots you conceivably might want to fill with Nvidia cards, but I wouldn't count on it. (FWIW, my company maintains the macOS driver for one of 2(?) manufacturers of USB graphics adapter chips, so I have a pretty good idea of how the macOS graphics stack works.)
- dis-sys 7y agoFor those who might be interested for a cheaper alternative - Intel Xeon 8280 QS is $1,800 USD each, a pair of those (56 cores) on a supermicro mb with 12 x 16G RAM is about $5k USD. It has an impressive Cinebench R15 score of 7,800+
- fabian2k 7y agoWhere do you get the $1,800 price? Intel says the recommended price is around $10,000: https://ark.intel.com/content/www/us/en/ark/products/192478/intel-xeon-platinum-8280-processor-38-5m-cache-2-70-ghz.html https://ark.intel.com/content/www/us/en/ark/products/192478/...
- dis-sys 7y agoI clearly mentioned it is an XEON 8280 _QS_. You can buy from ebay.
- fabian2k 7y agoLooking this up it seems like Engineering and Qualification samples are loaned by Intel, so the people selling them don't actually have the legal right to do that. They also might have other issues and limitations because they're not the final product, and you don't have any warranty on them. That's a lot of caveats to hide behind two letters.
- justinjlynn 7y agoIndeed. Stay away from Qualification Samples... the last thing anyone should want is an Intel CPU that's even less half-baked than usual.
- chx 7y agoOP said "QS" these are grey market chips, QS means Qualification Sample, Intel sends them to various manufacturers to test with motherboards and then they leak into the grey market. They are not retail so frequency and other characteristics might differ -- but even their stability might not be as high as a retail one. Although QS is usually better at this, it's ES (Engineering Sample) which can be very dicey.
- kjTAB 7y agoI don't believe 64 core, 3.35GHz at 200 Watt. Previous numbers for 64 core were around 2.2GHz. Looks like clickbait to me.
- huntie 7y ago3.35GHz is max boost frequency, they don't list the base clock speed. Last gen was 2.7-3.2 for the high-end chip.
- bitminer 7y agoThey offering a very wide range of core counts but only a narrow range of watts per socket. Question: for single threaded performance (*n where threads are independent) do I want higher watts per core or higher number of cores? By "performance" I mean throughput.
- piinbinary 7y agoI'm curious why they aren't charging more for it. Wouldn't that be free profit margin? Or perhaps they are trying to hurt Intel by selling it at a price where Intel can't make a profit if they lower their prices to be competitive? Or, maybe it has to do with trying to quickly grab market share. Maybe AWS and other purchases of servers have a somewhat fixed budget, so cheaper chips translates directly to more chips sold? Both those explanations seem unlikely to me.
- tmd83 7y agoOnly AMD execs truly knows but I think there might be two factors involved. They really want to attract new customers and build echo system and the second part is hard when everyone has been buying almost entirely intel. The second is their cost for such a cheap should be enormously cheaper than intel because of chiplet design. A single chip die for the intel chip is way bigger and yield and cost dramatically increases as die size gets bigger from what I read. So at same profit margin AMD chips would be cheaper and to gain the market share they are probably willing to lower their profit too so those two element adds up.
- metildaa 7y agoNote that this chip cost increase is due to silicon defects, a larger chip has more defects, thus your failure rate significantly increases with each increase in chip area. AMD is working around this by using 8 separate CPU chips, wired together with one interconnect chip.
- cududa 7y agoYour ecosystem argument really holds a lot of water with me. AMD said they wouldn’t be introducing a new socket type until 2021
- api 7y agoIntel is the standard and that has amazing sticking power. AMD is trying to make an offer DC operators can't refuse to break that. Intel is reeling from major design flaws and process stagnation right now. It makes sense for AMD to punch hard. Now is the time. Also note that the eternally predicted ARM64 wave into servers, workstations, and cloud is not materializing. So far nobody has been willing to build such high performance chips and price them aggressively enough. All things considered it's an amazing window for AMD to take the market lead.
- xvf22 7y agoFinding mainstream vendor offerings has been a bit more difficult with AMD. We recently bought 2 E-2146G servers from Dell (preferred vendor at the org.) because anything AMD was only offered as a 2 processor server which was overkill for the application. I would have gladly purchased AMD based servers instead if they were priced closely.