4 ms·
This is due to it essentially having 6 channel memory versus 2 channel memory for AMD/Intel? I would definitely like to see if the x86 industry could figure ou
by Osiris 3y ago
This is due to it essentially having 6 channel memory versus 2 channel memory for AMD/Intel?
I would definitely like to see if the x86 industry could figure out how to include more channels while also not requiring 6-8 DIMMS to take advantage of it, like Thread Ripper.
Something like dual channel on a single DIMM?
- bryanlarsen 3y agoI imagine they'll get it the same way Apple has done it with M series and AMD has done with MI300; by putting the RAM in the same package as the CPU. Doing so will break AMD's promise of 3 years of compatibility for AM5, so the change will either happen with Zen6 or they'll release a new range of chips called something other than Ryzen and sell Zen5 versions of both Ryzen and this new range. Intel & AMD are leaving too much performance and profit on the table not to do so eventually; IMO it's more a matter of when rather than if.
- jauntywundrkind 3y agoIt'll be interesting to see where this happens. I have to hand it to Apple, it's bold that they have on-package RAM that's so wide. (And then doubling and quadrupling whole core complex with tiling is a stunning move; Pro and Max.) Apple shipped a huge range at all once: a very capable mobile chip to a very beastly workstation grade chip. Intel's most notable attempt to me was Lakefield (2019), a Mobile Internet Device (MID) class (sub-laptop) chip that I quite liked. The 1+4 architecture wasnt very fast and it was a bit more power intense than the Snapdragons of the day that it had a modest-to-significsnt lead over. But I loved that it was like this tiny tiny package that you basically just had to add power to, and the on-package LPDDR4X-4266 was a very speedy offering for 2019. Intel's upcoming Lunar Lake mobile was shown in January, rocking tiled (doesn't look like a stacked/3D Foveros setup) on package ram. https://www.anandtech.com/show/21219/ces-2024-intel-briefly-shows-lunar-lake-chip-nextgen-mobile-cpu-uses-onpackage-memory https://www.anandtech.com/show/21219/ces-2024-intel-briefly-... But both AMD and Intel already have big on-package ram cores. Intel's been shipping a "Max" Sapphire Rapids with 64GB HBM2e for a year (https://www.tomshardware.com/news/intel-launches-sapphire-rapids-fourth-gen-xeon-cpus-and-ponte-vecchio-max-gpu-series https://www.tomshardware.com/news/intel-launches-sapphire-ra...). Like Apple's Ultra, it's a quad-tile with each CPU having a it's own HBMe stack. AMD's MI300A is basically a GPU where some of the tiles are instead CPUs but it too has 8 stacks of HBM3. I keep asking myself how & when & where is on-package ram going to arrive in. But there's already significant HBM presence in big cores! It didn't seem to make an huge difference for Sapphire Rapids; some help but unless one uses the expensive accelerators well there s probably not enough core to use it. Meanwhile MI300 is only just happening & more API focused. We have yet for on package ram to really be meaningful & available & making a difference like it did for Apple. But as your post says, it feels like an inevitability. Someone's gonna make a core that can do more or be better by having lower powered faster local ram. Part of my suspicion is that the market is resisting de-segmentation. The real issue is that Apple used on package ram to add many channels. Not of slow wise HBM memory, but mamy channels of DDR ram. These companies don't actually want to compete on throughout; they want throughput to be associated with $10k exotic chips. They are lament to build higher bandwidth more-channel consumer cores. That starts to change some next year with AMD's Strix Halo, a big APU with quad-channel ram. There's no on-package ram as far as I've heard, but once you start having that much board real estate & energy going to ram, it sure would be nice to get even more performance for less power & much less space. Damn I love on package setups. 2024 doesn't seem likely to offer much new or exciting, but 2025 has some possible signs for hope.
- paulmd 3y ago> I would definitely like to see if the x86 industry could figure out how to include more channels while also not requiring 6-8 DIMMS to take advantage of it, like Thread Ripper. you can't really do that, DDR5 has a concept called "pseudo-channels" where a normal 64b channel can be broken down into 2x32 smaller ones, which improves parallel efficiency somewhat (now you can have 2 requests in-flight at the same time). But mostly bandwidth is down to the number of pins and how much data you can physically push down them, chopping the same pins into smaller channels doesn't help, other than letting you eliminate some inefficiency/overhead. however, this is essentially the goal behind strix point/strix halo - narrower memory buses (compared to apple) with cache to improve the effective bandwidth. Just like in GPUs, this allows you to use fewer channels to get to the same bandwidth, which means less actual data movement. The downside is, of course, less "raw" bandwidth, if your workload is not cacheable. basically my rough expectation is that it's going to be more expensive than apple silicon, and still probably not actually beat on power, but will allow you to do workstation laptops with 512GB or 1TB of unified memory, which is also something the apple stuff cannot do. They are different products, apple is targeting people who want a powerful ultrabook, amd is targeting people who want a mobile threadripper for actual work tasks (metrology is one example).