12 ms·
AMD's Future in Servers: New 7000-Series CPUs Launched and EPYC Analysis
- hyperbovine 9y agoSoooo the Linux kernel now compiles in 15.6 seconds. Jeebus I feel old...
- sp332 9y agoI was about to point out that it compiled in 4.8 seconds back in 2002 http://es.tldp.org/Presentaciones/200211hispalinux/blanchard/talk_2.html http://es.tldp.org/Presentaciones/200211hispalinux/blanchard... But then I remembered that it was only 4 million lines of code back then (v2.5.x) and now it's 18 million! https://www.linuxcounter.net/statistics/kernel https://www.linuxcounter.net/statistics/kernel (Edit: fixed second link)
- hyperbovine 9y agoI haven't compiled it since the days of Slackware but it took at least on hour on my Pentium 90. (Then I discovered Debian.)
- claudiug 9y agoRemember gentoo Linux
- quickben 9y agoActually considering it for the next build. 16 amd core CPU. Probably 4 machines to crunch on something in the next few years. These are awesome times we live in.
- peller 9y agoI'll admit the endless compiling was often torturous, but I have the Gentoo docs to thank for most of what I know about the innards of Linux and operating systems. Bet I'm not alone in judging the time spent worth it :)
- thomasmeeks 9y agoMy time in Gentoo was absolutely worth it for the same reasons. There are dozens of us! Dozens! I still recommend it to people who are serious about diving into Linux.
- c3833174 9y agoI'd say I learned more by playing around with Slackware, but Gentoo provides all those nice tools to mantain a consistent system.
- stingraycharles 9y agoThere used to be a long running joke (I think perhaps Linus even coined it?) that the kernel codebase grew about as fast as CPU performance improved. I guess that died around the time AMD released their dual core CPUs..
- ethbro 9y agoIt seems reasonable, considering anyone working on it enough to need to recompile would have a minimum change time correlated with compilation time. (Incremental compilation aside)
- seanp2k2 9y agoGoing to get a cup of coffee takes the same amount of time as it did 15 years ago, so devs still need a compile time which is roughly equivalent. 15 seconds would be much too fast, for example.
- amiga-workbench 9y ago"What Intel giveth, Microsoft taketh away"
- derefr 9y agoIs that an apples-to-apples comparison (either both monolith or both all-modules kernels, both with all drivers enabled?) I recall, back in the long long ago, running through `make menuconfig` and disabling what I didn't need to get a smaller kernel + shorter build-times.
- sp332 9y agoNo, I'm sure lots of things have changed. More of the code would be using C99 semantics now, and I have no idea how GCC would have improved or regressed since then. It's possible that it's trading improved runtime performance on ever-more complex hardware for longer compile times.
- zbobet2012 9y agoThe kernel is still c89.
- sp332 9y agoI see that gnu89 is the default, to avoid compatibility problems with older code, but does that mean there is no code in the kernel which requires gnu99? I saw some work was done to make the whole kernel gnu11 compatible as well.
- fahadkhan 9y agohttps://stackoverflow.com/questions/20600497/which-c-version-is-used-in-the-linux-kernel https://stackoverflow.com/questions/20600497/which-c-version...
- sp332 9y agoRight, but I was too lazy to do my own analysis and just extrapolated from that 2013 post to guess that the number is higher now.
- mrb 9y agoThe difference being of course that this 4.8-second compile in 2002 was done on an IBM p690 which cost half a million dollars. Whereas the 15.6-second compile was done on the EPYC 7601 which is only $4k for the CPU, ~$6k for a whole machine. $6k vs $500k... Let that sink in :)
- dis-sys 9y agodid they fix the random crash issue when running gcc? https://www.phoronix.com/scan.php?page=news_item&px=Ryzen-Compiler-Issues https://www.phoronix.com/scan.php?page=news_item&px=Ryzen-Co...
- pulse7 9y agoAMD Forum Thread: https://community.amd.com/thread/215773 https://community.amd.com/thread/215773 Last response from AMD: "The vast majority of users using Ryzen for Linux code and development have reported very positive results. ... A small number of users have reported some isolated issues and conflicting observations. AMD is working with users individually to understand and resolve these issues."
- 0xcde4c3db 9y agoI doubt there's any single root cause here. "I get crashes when running big compiles" is a classic symptom of hardware that is almost, but not quite, stable. This can be due to power delivery problems, thermals, marginal RAM, overclocks, faulty CPU (yeah, it happens sometimes) and so on.
- agumonkey 9y agoRun tccboot iso, you'll have your 15minutes of youthful fun https://bellard.org/tcc/tccboot.html https://bellard.org/tcc/tccboot.html
- setq 9y agoI still remember doing a make world on FreeBSD on a 75MHz Pentium. Went out for the night and it was still going when the hangover wore off the following afternoon. This was mainly because it spent most of it's time swapping.
- Symmetry 9y agoThe browser is the new OS and the compile times prove it.
- ekianjo 9y agosome of the reduction is not just hardware related if i remember correctly. Compilers have improved as well.
- DuskStar 9y ago4 dies per package is a pretty interesting way of doing things - probably helps yields immensely, but I can't imagine it does anything good for intra-processor latency. 142 ns to ping a thread on a different CCX within a die isn't too horrible, but I really want to know what sort of penalty you'll have from going to a different die within a package.
- dom0 9y agoMCMs aren't exactly new, especially not for bigger iron, though.
- greglindahl 9y agoNo, and there is a whole zoo of them shipping now. Xeon Phi has Omni-Path on-package, for example, it's a separate chip.
- sliken 9y agoI think it's less about yield, than it is about amortizing a large R&D budget across more chips. Keep in mind that they are competing with Intel who ships huge volumes and is using this generations profit to fund the next generation fabs. So AMD pours everything into a single die and manages to hit the desktop (Ryzen with 1 die), workstation (thread ripper with 2 dies), and server (Epyc with 4 dies). All with a single wafer of silicon, and as a bonus the memory bandwidth, max memory capacity, and total number of cores scales to fit all 3 markets. Pretty crafty. This is far from new of course, various Power CPUs, Intel cpus as far back as the pentium pro, and of course the previous generation Xeons. Intel's strategy has been one die for most of their low end/low power chips (max 2 core/4T) and a larger die for their desktop (max 4c/8t) that use the same socket (LGA1151). Then Intel targets HEDT (High end desktop) and server with the same chip, same die, same socket, just marketing to differentiate the x-servies chips and the regular single/dual socket chips that share chipsets, sockets, and a lga2011 socket. AMD seems to have scared Intel pretty bad. They have held off on the skylake xeons, only shipping them to cloud providers while they wait for the AMD release before releasing skylake xeons to the masses. I'm glad to say AMD performance seems pretty good, SpecINT (using GCC) is around 50% faster than the similar Intel chip and SpecFP is even better. Seems fair unless you use the Intel compiler to compile all your binaries anyways. In fact the fastest AMD + gcc-6.2 is faster at SpecFP than the fastest Intel + Intel's compiler (1330 vs 1090 respectively).
- satai 9y ago1 socket 16 / 32 @ 2.9GHz max for $700+... it looks like 16 core Threadripper with reasonable frequencies for less then $999 looks in reach...
- brianwawok 9y agoSo how will threadripper different from the single socket guy presented here?
- snovv_crash 9y agoThreadripper is a 2-chip module rather than 4, so will have half the L3 cache. Threadripper has half the memory channels and half the PCIe lanes. EPYC is available with up to twice the cores. The main advantage I see for Threadripper is that at 16 cores EPYC will have half the cores disabled, so for problems that fit in L2 you lose some performance from the reduced L2 sharing. That and it should be priced better than server chips, with I suspect the 12 - 14 core being the sweet spot. Based on leaks I think Threadripper might boost a few 100 MHz higher, with base clock up to 3.6GHz.
- jnordwick 9y agoIs L2 shared on AMD or are they going the Intel route and making it per core? Is there any info on cache architecture for Zen?
- dom0 9y agoHm? L2 is per core. Always has been in a three+ layer architecture. Zen has 512K L2 per core and 8M L3 per CCX (two CCX per dice). L3 is a victim cache iirc, unlike previous generations where the L3 was inclusive. Intel usually went with a similar scheme in the last few years, where the L3 is partitioned into slices assigned to cores; accessing the local slice is faster than a non-local slice. Skylake-SP deviates from this (significantly), for better ... or worse.
- 9y ago
- nik736 9y agoWhy does AMD compare their single socket CPUs to Intels E5-2XXX line? Intel has E5-1XXX single socket CPUs.
- satai 9y agoAll the E5-1xxx v3s, v4s are 8 cores or less. That's probably the reason.
- Terribledactyl 9y agoE5-1XXX core count is much lower, only 8 in v4 (latest shipping)
- sp332 9y agoI think the idea is that single-socket EPYC CPUs beat many Intel dual-processor setups. If you break it down by sales numbers, a single EPYC might beat the most popular dual-Intel platforms. http://images.anandtech.com/doci/11551/epyc_tech_day_first_session_for_press_and_analysts_06_19_2017-page-022.jpg http://images.anandtech.com/doci/11551/epyc_tech_day_first_s...
- nik736 9y agoOh wow! I totally overlooked the 2x E5-2XXX part. That makes much more sense now and is totally awesome.
- sliken 9y agoThe lower end dual socket E5-2xxx are the most popular servers. AMD claims they can win on performance, match or beat maximum memory, and win on perf/watt when comparing a single socket AMD to a lower end dual socket Intel. Seems reasonable to me, I'd consider getting a quad system in 2U amd system with single sockets if it beat price and perf of a quad system in 2U intel system with dual sockets.
- garaetjjte 9y ago>In this case, an EPYC 7281 in single socket mode is listed as having +63% performance (in SPECint) over a dual socket E5-2609v4 system. So, quad-CPU is faster than dual-CPU? Not surprising.
- mtgx 9y agoI guess the point is AMD offers significantly more bang per buck (and socket) at similar prices.
- dragontamer 9y agoThe per-socket performance might be a nice "marketing hack", since a lot of big-iron software is sold and licensed as "per-socket". IIRC, Windows Server is sold per-core however. So lots of cores may rise the total-cost of ownership in the case of Windows Server.
- deleted 9y ago[deleted]
- iamtherhino 9y agoMost enterprise database vendors now charge on a "per-core" basis. There are some OLAP / NoSQL vendors that are moving to a per XX GB Ram model.
- maksimum 9y agoHow are these vendors getting away with this? Is their product/support so much better than open source + third party support, or are they entrenched?
- _ix 9y agoInertia and vendor lock-in are really powerful forces. Previous decision makers chose a Microsoft stack, and those who came before them did the same. Our third party services provider was a staunch Microsoft supporter. Before anyone knew it, we were completely locked in to a Microsoft ecosystem, paying tens of thousands per core for new SQL Server licenses.
- gbrown_ 9y agoThose TDPs look pretty high, what are vendors willing to put into 1U high 0.5U wide style servers with 2 sockets these days? Last I looked I seem to recall it was around up to 145W.
- quickben 9y agoAll of them? Nothing Intel made so far can compete with the efficiency.
- lettergram 9y agoI would agree, but there appears to be a 15% to 20% improvement of efficiency. Meaning, the processing power per watt is worth more.
- sliken 9y agoKeep in mind that AMD puts more on chip than Intel, so you should look at system power, not socket power. The comparisons so far I've seen show AMD doing as well or better than Intel on perf/watt.
- satai 9y agoAMD EPYC 7601 Dual Socket Early Power Consumption Observations https://news.ycombinator.com/item?id=14598660 https://news.ycombinator.com/item?id=14598660 https://www.servethehome.com/amd-epyc-7601-dual-socket-early-power-consumption-observations/ https://www.servethehome.com/amd-epyc-7601-dual-socket-early...
- rb808 9y ago> Running an AVX2 workload we were expecting much higher power consumption but at under 500w for 128 threads, this is excellent. eish
- revelation 9y agoGranted the AVX2 performance of the Zen processors is intentionally crippled. It's more of a software compatibility implementation than actual performance boost.
- girst 9y agointel had a monopoly on high-end chipsets for _far_ too long. I'm glad, there is some competition.
- irishjohnnie 9y agoWow! AMD EPYC + Xilinx FPGA!
- digitalzombie 9y agoI'm... actually shock that somebody care about Xilinx. I had to do Xilinx in my CE classes. It was terrible software, we joked that the CE and EE people code that software. Crashed all the time, made me so paranoid that to this day I would often ctrl+s every few minutes just in case my IDE crashes.
- q3k 9y agoFWIW new Xilinx silicon is programmed by a new software suite, Vivado - which is much better than the terrible ISE you probably had to use in your class.
- myrandomcomment 9y agoI would really love it if there was a benchmark around running VMs and containers for something like this. Our dev/test system is all docker containers so that is what we would care about. I guess it would be hard as there are to many ways to scale out what you run - how many VMs, how many containers, what are you running in them? It would be an interesting benchmark matrix to sort for. It would be interesting just to see how many containers you could start, run lighttpd and each server a static web page? Maybe 1/2 with the page and 1/2 with an application that builds the page? Who knows...to many variables. I think we will just by a system when we can and try our workload on it. Oh, well.
- chrisseaton 9y agoI don't think running processes inside a container will be any more interesting for benchmarking than just running the processes normally. What overhead does a container add? A little indirection in syscalls? I would imagine the number of instructions involved in that are beyond trivial for serving a page, so I can't see how benchmarking running containers would be different than just benchmarking running your processes. VMs - now with processor virtualisation technology I'm sure the different processor architectures do make an interesting difference there.
- myrandomcomment 9y agoIn some of our scaling test where each container had an IP connecting out to a remote system we ran in to a ton of issues at scale that we did not see running the same number of processes on the bare metal. The overhead of the namespace can add up.
- nwmcsween 9y agoThis is due to the what the container orchestratior uses to route traffic, check out project calico for low overhead networking.
- myrandomcomment 9y ago
- dang 9y agoRelated: https://news.ycombinator.com/item?id=14598660 https://news.ycombinator.com/item?id=14598660
- jnordwick 9y agoI didn't see any info on the cpu cache architecture which governs performance for many applications now. Anybody have any info on things like L0 to L2 size, type, latencies, etc?
- choudanu4 9y agoThey mentioned in the article that the full cache of each die is available. Additionally, EPYC uses the same dies used in Ryzen. I'd look at earlier articles for Ryzen to determine latencies within a single die. So for whatever cores are enabled on each die, you get the L1/L2 caches for each core as per the Ryzen launch. Additionally, you get all of the shared L3 cache, irrespective of the number of cores disabled per core complex. This pattern follows across all four dies in each socket.
- jnordwick 9y agoI read that, but how much, type, architecture, latencies, etc? This is a huge factor in the performance of the chip.
- tyingq 9y ago"Each Epyc has 64KB and 32KB of L1 instruction and data cache, respectively, versus 32KB for both in the Broadwell family, and 512KB of L2 cache versus 256KB. AMD says Epyc matches the Broadwells in L2 and L2 TLB latencies, and has roughly half the L3 latency of Intel's counterparts." https://www.theregister.co.uk/2017/06/20/amd_epyc_launch/ https://www.theregister.co.uk/2017/06/20/amd_epyc_launch/ Various L3 sizes in the article.
- IanCutress 9y agoOur original Zen Microarchitecture deep dive has all the info. http://www.anandtech.com/show/11170/the-amd-zen-and-ryzen-7-review-a-deep-dive-on-1800x-1700x-and-1700 http://www.anandtech.com/show/11170/the-amd-zen-and-ryzen-7-...
- 9y ago
- bsaul 9y agoA bit off topic, but does anyone knows if AI ( aka modern neural networks) plays a role in cpu design nowadays ?
- m-j-fox 9y agoHere's one CPU that has nn in mind: https://cloudplatform.googleblog.com/2017/04/quantifying-the-performance-of-the-TPU-our-first-machine-learning-chip.html?m=1 https://cloudplatform.googleblog.com/2017/04/quantifying-the...
- buryat 9y agonot sure if they play a role in designing CPUs, but AMD says that their Zen CPUs have a neural network inside for branch prediction [1] which is not something new [2] [3] [1] https://www.amd.com/en/technologies/sense-mi https://www.amd.com/en/technologies/sense-mi [2] https://www.jilp.org/cbp2014/paper/DanielJimenez.pdf https://www.jilp.org/cbp2014/paper/DanielJimenez.pdf [3] http://cseweb.ucsd.edu/~atsmith/rnn_branch.pdf http://cseweb.ucsd.edu/~atsmith/rnn_branch.pdf
- irishjohnnie 9y agoProbably not. Do you use AI to write code?
- Symmetry 9y agoA compiler does a lot of AIish things in turning my C code into a sequence of instructions. Much more classical AI than machine learning but still within the AI umbrella.
- irishjohnnie 9y agoI was referring to the RTL-based IP that goes into each subsystem of the CPU, hence the analogy to code. You're talking about compiler instruction scheduling, a lot of which are a bunch of non-AI algorithms. If I'm missing something, I would appreciate the references to the functionality you're referring to.
- 9y ago
- Keyframe 9y agoWhat's the SSEs and AVXs performance like on Ryzen/EPYC compared to intel?
- gcp 9y agoSSE performance is comparable, so is 128-bit AVX. 256-bit AVX is in theory half as fast. Despite this, apparently Ryzen beats Kaby Labe on SPECfp, so theorethical max throughput is only part of the story. It does not help that Intel needs to heavily reduce their boost speeds when using the full AVX unit.
- TazeTSchnitzel 9y agoI wonder how the AVX situation will play out. Will 128-bit SIMD with more cores and higher clockspeed beat 256-bit?
- AlphaSite 9y agoIt's not actually half as fast, because these chips run a AVX at the full clock speed, rather than down clocking as Intel does.
- Keyframe 9y agoSo, any benchmarks regarding SSE and AVX out there yet?
- oakridge 9y agoFor the new AMD releases there's none yet, but for Ryzen 7 Phoronix did one comparison for code compiled with -mavx2 and Ryzen did really poorly. The other benchmarks which use more integer math or relies heavily in multithreading puts AMD in an advantage. The post was in 18 May 2017 and AMD may have released some microcode updates that nullifies the results. The review: http://www.phoronix.com/scan.php?page=article&item=ryzen-kabylake-may&num=6 http://www.phoronix.com/scan.php?page=article&item=ryzen-kab...
- mastazi 9y agoFor anyone looking for info about the socket: * Epyc uses socket SP3 https://en.wikipedia.org/wiki/Socket_SP3 https://en.wikipedia.org/wiki/Socket_SP3 * Threadripper uses socket TR4 https://en.wikipedia.org/wiki/Socket_TR4 https://en.wikipedia.org/wiki/Socket_TR4 * Sockets SP3 and TR4 have the same number of pins (4094 pins) and they have the same cooler bracket mount (see https://www.overclock3d.net/news/cases_cooling/noctua_showcase_epyc_threadripper_ready_tr4_sp3_ready_cpu_coolers/1 https://www.overclock3d.net/news/cases_cooling/noctua_showca... ) * However they are still two separate sockets so you shouldn't expect to be able to use Epyc on TR4 or Threadripper on SP3
- Kubuxu 9y agoIt is probably to differentiate motherboards. You don't want the consumer to plug in Epyc into Threadripper's mobo and complain that there is not enough PCI-E lanes or memory channels.
- dbcooper 9y agoBaidu and Microsoft will be customers: https://www.bloomberg.com/news/articles/2017-06-20/amd-server-chip-revival-effort-enlists-some-big-friends https://www.bloomberg.com/news/articles/2017-06-20/amd-serve...
- equasar 9y agoHP and Dell announced their new line of Servers based on EPYC. According to the keynote, thet were working with AMD since day 1 of EPYC development.
- greptomania 9y agoWhile I'm excited to see AMD's offering, as a scientific-HPC user I can't help but wonder how much marketshare AMD will be able to gain without more information on supporting software - specifically good compilers + math libraries (cf. Intel compilers + MKL). Strangely, I've not seen much on HN, or elsewhere, make mention of AMD's software support. Is this because it doesn't exist, or because compilers are less "sexy" than shiny new hardware?
- AlphaSite 9y agoThey have a custom fork of Clang http://developer.amd.com/amd-optimizing-cc-compiler-aocc-technical-support/ http://developer.amd.com/amd-optimizing-cc-compiler-aocc-tec... for C and ROCm for GPUs https://github.com/RadeonOpenCompute/ROCm https://github.com/RadeonOpenCompute/ROCm
- geezerjay 9y ago> While I'm excited to see AMD's offering, as a scientific-HPC user I can't help but wonder how much marketshare AMD will be able to gain without more information on supporting software - specifically good compilers + math libraries (cf. Intel compilers + MKL). My take is that AMD's newest offering will be very well received by everyone who has a relatively tight budget but needs a small supercomputer on the desktop. This means data analysts and people doing all kinds of structural analysis work. As some optimization algorithms fit the definition of embarrassingly parallel, the expected turn-around time of anyone doing that sort of work will benefit greatly from the extra speed, bandwidth and core count of AMD's Ryzen/Threadripper/Epyc line.