7 ms·
For days, I'm reading about the 888, its Cortex X1 and a 35% uplift and that it's coming (some day) but I couldn't find a single word about how it will perform
by desmap 6y ago
For days, I'm reading about the 888, its Cortex X1 and a 35% uplift and that it's coming (some day) but I couldn't find a single word about how it will perform against the M1. Even Anandtech dares to bring four, full pages of fluff.
This is what most are interested in and Qualcomm should asap disclose the performance and/or give an outlook on what is coming, always referencing M1's performance as the current benchmark.
- baybal2 6y ago888 will likely be still n-times slower than Apple M1 per-core/clock, even if you fit it with equally fast I/O, and cache. I believe X1 will still be able to beat Firestorm on size/performance ratio.
- desmap 6y agoThis. Exactly my gut feeling and the press is not questioning anything just rephrasing press releases from Qualcomm, writing about so important details like why the name 888 and so on. Qualcomm and Intel should come out asap and tell what they plan to do in the next 12-18 months, e.g. embedding RAM or whatever, to match M1' performance. But they rather work on original naming schemes. Good that I sold my Qualcomm stocks years ago.
- SSLy 6y ago>embedding RAM or whatever Due to the different way the cells are made that'd be very, very expensive.
- Demiurge 6y agoCan you explain why 888 will be n-times slower? Assuming, with n greater than 1. I just looked up the Snapdragon performance vs Apple A14, and it is within ~10% [1]. I have also compared it to an intel processor, and it also ahead [2]. From a quick search, it looks like Qualcomm is a competitive contender, if we're going to ARM. If true, that would be a pretty exciting new world of ARM PCs :) [1] https://nanoreview.net/en/soc-compare/qualcomm-snapdragon-865-plus-vs-apple-a14-bionic https://nanoreview.net/en/soc-compare/qualcomm-snapdragon-86... [2] https://gadgetversus.com/processor/qualcomm-sm8250-snapdragon-865-vs-intel-core-i7-8650u/#:~:text=The%20processor%20Qualcomm%20SM8250%20Snapdragon,Snapdragon%20865%20was%20designed%20earlier https://gadgetversus.com/processor/qualcomm-sm8250-snapdrago....
- desmap 6y ago> it is within ~10% This is just so off. Look at Geekbench's figures for a 865 Plus, add 35% on top and you are still not close to an A14 and miles away from the M1.
- Demiurge 6y agoOk, I'm sorry, I haven't been reading CPU reviews for the past... 10 years. Can you explain why I would add 35% on top? What is "miles away" converted to percent difference in operations per second?
- Leherenn 6y agoI think the 35% on top is the expected improvements of the 888 compared to the 865. On Geekbench, the top 865 is at 887 for single thread compared to 1585 for the A12. If we add 35% to the 865, the expected 888 score is ~1200, still quite behind the A12.
- desmap 6y agoThanks for doing the math.
- Demiurge 6y agoI see, got it. So, if we look at single core performance, I can see it is much less performant, but does not seem like more than a magnitude. It would also make it more efficient than Intel. Still, overall, would be very interesting to see a high performance Snapdragon ARM laptop!
- zipityzi 6y agoActually, Qualcomm noted the 888 is just 25% faster than the 865 in "overall performance". And that's Qualcomm's marketing number. So Qualcomm didn't even compare the 888 to their latest 865+ CPU: it'll be less than 25%, but how much? Source: https://www.anandtech.com/show/16271/qualcomm-snapdragon-888-deep-dive https://www.anandtech.com/show/16271/qualcomm-snapdragon-888... The 25% number is compared to the 865, not the 865+
- skavi 6y agoThis doesn’t compete in the same space as the M1. It only sort of competes with the A14. And it will be quite a bit worse performing than that.
- desmap 6y agoSince the 855 or even earlier, Qualcomm had always a notebook derivate SoC (which were not much faster than the mobile phone version). However, the M1 is now the reference for any ARM licensee and anyone at Qualcomm's scale should have some answer to Apple's SoC, especially then when the create some press buzz around a new SoC with a fancy name.
- piyh 6y agoThe magic of M1 is its x86 emulation speed with it's memory flag which I'm guessing is not something that has been on the roadmap for Qualcomm until recently. Even then you need integrate all the hardware/software partners instead of Apple's vertical integration. Even if they have M1 level performance without the good emulation, what's the point if it's stuck with windows RT?
- desmap 6y agoWindows RT isn't too bad. You wouldn't actually notice a difference. A decade ago (or less), I had the first RT-based Surface running native Word processing a 300-pages PhD paper, it was running buttersmooth and cool, all the time. I think this is more about MS' commitment to ARM than a software problem.
- mattnewton 6y agoWeird, my experience with that same device when it launched was noticible input lag while trying to use explorer and microsoft word at the same time. The stuttering was so bad, scrolling and typing could take 100's of ms - now I wonder if I wrote off the whole RT experiment because I got a lemon or malware.
- Nokinside 6y agoWhile Apple did really stellar job with M1, Apple's R&D counts less than 50% of M1 performance gain over competition. Rest is TSMC's 5nm and perfect timing. Apple generates largest revenue per processor by a wide margin. This means they can afford to pay premium to be first at everything and have 6-12 month lead. TSMC's 5nm volume production and Apple's 2020 launch was timed to go together since 2017 (if go back to TSMC conference call you can hear them talking about it without mentioning details).
- whynotminot 6y agoA lot of the lead is process, no doubt, but a lot of it is simply a better design. Every step of the way, even when they're at process parity, Apple chips have been better.
- Nokinside 6y ago50% process technology. 25% for the fact that Apple can choose more expensive solutions in every turn. Better performance is often more expensive. 15% because Apple can design only for Apple products. Qualcomm, Intel, AMD .. have to design more generic products. ~10% at most for something else.
- whynotminot 6y agoYour math doesn't make sense. 25% because Apple chooses better solutions, and then you finish with only 10% better design? Is it only a better design if it's not more expensive? So you get to decide that? Does BMW not get credit for better design than a Ford Focus because it's more expensive? I guess it doesn't really matter. Apple makes a better product. That's kind of the story here at the end of the day.
- zipityzi 6y agoWhat are these numbers based on? Just your "best estimate"? It's dismissive of Apple's uarch for what reason? Arm has architecture licenses, unlike x86: anyone could've designed an Arm CPU from the ground up. NVIDIA tried, AMD tried, Intel tried, Qualcomm tried, Samsung tried, Huawei tried, etc. Everyone had a chance (and they still do). Arm is the most level playing field available in high-perf CPU design. And "50% process technology" is an exaggeration even embarrassing for HN. The A13 was built on 7nm and still beat perf/watt of any x86 CPU and total 1T performance rivals Tiger Lake. A14 / M1 are a natural evolution of that same uarch. https://images.anandtech.com/doci/16226/perf-trajectory.png https://images.anandtech.com/doci/16226/perf-trajectory.png What really breaks down your argument: Samsung has had nearly every advantage as Apple, yet its Exynos line is some of the worst-perf/watt Arm uarch today: its own OS (Tizen), its own foundry, its own phones / tablets / laptops, and a massive conglomerate for funding. What happen to Samsung? What money doesn't Samsung have? Hell, Samsung is even MORE integrated than Apple, as Apple still needs to outsource its fabrication to TSMC. There's a reason Samsung is giving up on its uarch and moving to Arm stock cores (X-1, A78, A78C, etc.), just like Qualcomm + NVIDIA. https://news.samsung.com/global/samsung-electronics-announces-fourth-quarter-and-fy-2018-results https://news.samsung.com/global/samsung-electronics-announce... Don't tell me "you need trillion dollar valuation to make a top-class CPU". AMD was nearly bankrupt 6 years ago and now has the fastest general compute x86 arch in the world. This argument reeks of "well, if Apple has the fastest CPU uarch today, then I'm going to ensure everyone else has an excuse." Apple didn't even have an Arm architectural license, much less a custom high-perf CPU, 13 years ago. 13 years from "never designed a high-perf CPU" to "dominating x86 perf / watt while at the heels of total perf" is actually notable and actually impressive.
- bfgoodrich 6y agoSome Geekbench results have appeared and seem to show the 888 still at a pretty big disadvantage compared to the A14. e.g. the big cores in the A14 offering 40% greater performance, so much so that the 2 big+4 little of the A14 has higher MT performance than the 4 big+4 little 888. These results could very well be fake, however early in a product release OEMs do often intentionally release these teaser numbers. This would make it the fastest Android chipset by a good margin, and bring it closer to equalling the A12 of 2018.
- sradman 6y agoSeveral Big Tech vendors are designing ARM SoCs as outlined in the recent HN post The Tech Monopolies Go Vertical [1]. The Snapdragon 888 seems to compete with the Apple A14. I wonder why there isn’t more comparisons between the A14 iPad Air and the M1 MacBook Air. Microsoft’s SQ2 chip for the Surface integrates a better GPU into a Qualcomm based SoC. Perhaps GPU performance is a key differentiator. I find it odd that the 2+4 big.LITTLE A14 isn’t trounced by the 4+4 big.LITTLE Qualcomm SoCs. A dark horse, especially with respect to GPU, is Nvidia after their ARM acquisition. These are interesting times. [1] https://news.ycombinator.com/item?id=25251229 https://news.ycombinator.com/item?id=25251229
- officeplant 6y agoWhat I find the most interesting is Mediatek, who largely gets ignored among the major players. Their affordable SoC's powered the Lenovo Duet into being one of the most compelling affordable Chromebook Tablets of the year. They've also mentioned plans to put more powerful chips into chromebooks in 2021. ARM powered chromebooks run Android apps far better than their x86 cousins plus still include linux support (crostini) which has a growing set of ARM applications every year. I'm pretty excited for where the future of efficient computing is going. Between my M1 Mac Mini and overclocked Pi4 8GB I've managed to sell off and move on from x86 machines for personal use. Interesting times indeed.
- wyldfire 6y agoM1 competes with Qualcomm's 8cx / SQ2. It's certainly reasonable to show the comparison between M1 and 888 but the products that M1 show up in (Macbooks) don't generally go head-to-head with the products 888 shows up in.
- m463 6y agoOff-topic - where did "uplift" come from? This is a new word in the last 6 months or a year.
- hajile 6y agoA78 is basically a mild refinement to A77. X1 is a souped-up A78. * Move from 4 to 5-wide decoder (A14 is 8-wide) * Double number of 128-bit SIMD from 2 to 4 (A14 is 4) * Moving multiplication into a second ALU allows simultaneous multiplication and division (on both A78 and X1) * Minimum L1 cache at 64kb instead of 32kb with a 64kb option (A14 is 128 d-cache and 196 i-cache) * L1 bandwidth on A78 and X1 doubled over A77 (probable 2-4x as wide on A14) * Double L2 cache from 512kb to 1mb (A14 has 8mb L2, but no L3) * Double maximum amount of L3 cache from 4mb to 8mb (for a potential 9mb cache vs 8mb for A14) * Still ARMv8.2 (A14 is ARMv8.4 which brings some extra instructions for pointer security, virtualization, JS integer conversion, complex number SIMD, SHA512/SHA-3 hardware, int dot products, and some other stuff -- probably the biggest performance difference here will be the 1-2% for JS conversion) * Increase BTB from 64 to 96 entries (Not sure about A14, but probably more) * Increase TLB from 1k to 2k pages (A14 is 3072 pages) * mops bandwidth goes from 6 to 8 mops (not sure about A14, but 8+ for sure) * uop cache goes from 1.5k entries to 3k entries (I seem to remember hearing 2k entries for the A12, but I don't know) * mops dispatch also goes from 6 to 8 (once again , not sure about A14, but 8+ for sure) * uops dispatch increases from 10 in A77 to 12 in A78 to 16 in X1 (I'd guess Apple's at least this wide) * reorder buffer increasing from 160 to 224 entries (A14 ROB is somewhere around 630 entries) ARM's claims that with all three processors at 3GHz, the A78 is 7% faster than A77 and the X1 is 20-30% faster depending on the workload. I'd guess that SIMD float-heavy performance between X1 and A14 is very close. More interesting is the x86 comparison (let's look at zen 2). X1 has twice the L1 cache (64 vs 32k), but considering AMD lowered this amount, I doubt it's a real advantage more than different ISA and microarchitectures. Likewise X1 with 1mb of L2 vs 512kb for zen 2 probably doesn't matter too much (AMD redid the cache in zen 3 and stuck with 512kb while Intel still uses 256k). Zen 2 has 16 L0 BTB vs 96 L0 BTB for X1, but this might be microarchitecture more than anything as well as X1 drops to 2k L2 entries vs 7k for Zen 2. I don't know if 8mb L3 is shared in X1, but I'd assume so in which case 16mb for zen 2 is a definite advantage. The decode width is now 20% wider than zen, but x86 instructions are around 15% more dense, so I'd guess it's about the same. X1 has 3k uop cache entries vs 4k on x86, but I suspect x86 needs more to avoid hitting those complex decoders at all costs. AMD dispatches 6 int uops and 4 float uops per cycle while X1 does 16 uops per cycle. I don't know about ARM, but AMD's perceptron predictor is probably better (it seems like Apple or ARM would make a big deal about this if they'd made the investment -- Samsung did). Given AMD has 19 stages instead of 13 in the X1, a better predictor is needed to offset penalties. X1 and Zen 2 both have identical 224-entry reorder buffer. X1 has 2 AGU, 1 AGU/load and 2 load/store while zen 2 has 3 AGU, 2 load/store, and 1 store. Zen 2 and X1 AGU are both 256-bit wide. I can't find store queue length for X1, but bigger queues could be an advantage to AMD (I suspect so as ARM probably wouldn't want the extra power draw). A76 seems to have similar instruction cycle counts to Zen 2 and I assume they haven't gotten worse since. X1 has four 128-bit SIMD while Zen 2 has two 256-bit SIMD. If you can use wider SIMD, AMD has an advantage, but if there's lots of smaller SIMD, then X1 should have an advantage). Zen 2 has 4 integer ALU with one capable of multiplication, another capable of division, and another with a CRC (cyclic redundancy checking). X1 is identical (except I don't know if it has a CRC). X1 very likely comes close to Zen2 and even A14 in integer performance while being a bit more bottlenecked in float performance. With not one, but two ARM designs now taking on x86 so handily, perhaps it's about time we start considering that ISA maybe does matter more than Intel has claimed over the years.
- Causality1 6y agoWhy? Is Apple planning on releasing a phone running the M1?