4 ms·
> really crazy optimizations that they are doing Any examples of this?
by pixelesque 1mo ago
> really crazy optimizations that they are doing
Any examples of this?
- PunchyHamster 1mo agoThere aren't ones that wouldn't be served better by ARM The improvement is entirely "we don't have to pay ARM"
- zephen 1mo ago> There aren't ones that wouldn't be served better by ARM Sure there are. If you want to do weird and wacky stuff, try to do it with an ARM and see how fast you get shut down. > The improvement is entirely "we don't have to pay ARM" No, it's "we don't have to beg and grovel, or pay ARM."
- sylware 1mo agoI guess this would be orders of magnitude more horrible with x86_64?
- zephen 1mo agoHard to know at this point. In the past ARM sued people for daring to think about implementing ARM compatible computers. (Even university research projects.) Of course, in the more distant past, AMD and Intel sued each other a lot. But there have been a dozen or so x86 market entrants. I think the competition has just been brutal. So, with either one, you could, of course, do a clean-sheet design and hope you don't get sued. Since ARM has historically only sold IP, they are incredibly jealous of their monopoly for that ISA. Of course, the flip side of it is that if you have enough money, and you beg and grovel enough, you can probably do what you want with their code. Except, of course, when you can't. Or rather, maybe you can, but they will try really hard to stop you. For example, Qualcomm, an architecture licensee who had purchased the rights to basically make whatever the fuck ARM they wanted, purchased a startup named Nuvia that also had an architecture license, and started using Nuvia's designs. ARM sued, just because. They were shut down in court, but the message is clear. They are very aggressive. Seriously, who needs that shit? Semiconductor market windows are tight enough as it is, and ARM has always acted like "Nice chip you've got there, buddy; shame if you haven't dotted your i's anc crossed your t's on the licensing." At least with Intel or AMD, you're probably only looking at patent lawsuits, and you could probably do a decent job on a low end machine with techniques that were known 20 years ago. You could do the same with ARM, of course, but they will sue you just because, to see if they can make you run out of money.
- sylware 1mo agoAllright, you did describe a monstruous IP mine-field. No way anybody could reasonably use anything else than RISC-V.
- zephen 1mo agoBleeding edge processor development is always going to be a patent minefield, even for RISC-V. Most patents are broad enough to cover more than one architecture. But there are now several RISC-V IP vendors, and ARM can't possibly shut them all down, and companies can develop IP in-house, so there are lots of different designs already out there. So, at this point, developing a RISC-V processor provides a certain amount of safety in numbers, kind of like swimming in a school of fish, while developing an ARM compatible processor paints a target on your back. Developing a low-end x86 processor wouldn't attract any legal attention either, but now that the RISC-V ecosystem is big enough, why bother? Decoding all those instructions might be a lot of extra work for no real market share gain. Developing a high-end x86 processor would probably attract careful scrutiny of exactly how you managed to get that performance and whether you violated any patents to do that. Again, that the same sort of patent scrutiny would happen with bleeding edge RISC-V, but certainly, a fast x86 processor would be higher on the priority list of the Intel and AMD legal departments.
- sylware 1mo agoThen RISC-V would need something like they did for open source software (in countries with barbaric IP), some sort of pool of defensive IP. Because, if you look at the current trajectory of RISC-V micro-architectures: they will be competitive, so current hardware big tech will try hard to shut them down. If they are already at Zen3 performance... I run Zen2 and it's already beyond than enough (I miss AVX512 extensions though, hope we'll get often RVA cache line vector instructions on RISC-V). Look at qualcomm already buying what seems to be a very good RISC-V microarchitecture, wonder what they will do to it (usually, they let it die slowly forcing their good engineers to work on something else). The right way(TM) would be for ARM to "drop its ISA" and become a microarchitecture designer able to decode and run RISC-V ISA. Intel and AMD should too. The world would be basically competing micro-archs all with high compatibility at machine code level. And it seems nobody has been addressing the elephant in the room: x86 is an horrible ISA for a modern CPU, much more worse than RISC-V/ARM. It is only because AMD/Intel throw tons of money with tons of brains to make that horrible ISA performant AND because they are hogging the production capacity of the best silicon process. There is zero wonders here.
- camel-cdr 1mo agoHere are a few random things I know of: * Tenstorrent Ascalon has a neat optimization for certain LMUL>1 SIMD operations. LMUL=2 effectively unrolls the SIMD operation making it read two SIMD registers from every source and write two SIMD registers to the destination. There are however some instructions where LMUL=2 only needs to write to one registers, those are narrowing instructions (e.g. 64-bit to 32-bit truncation) and comparisons (which write to a LMUL=1 register with packed bits). When those SIMD instructions have to .vx form, which means one argument comes from a GPR, they now only need to write one SIMD register and need to read two SIMD registers. This matches what regular SIMD instructions need and because the silicon for the execution is much cheaper than register file ports, Ascalon can exexute these instructions in a single operation. So you can compare twice as many SIMD elements against a scalar, then you can against another SIMD register. * Ventana (now under Qualcomm) talked a tiny bit about their fetch-block-optimizer and something that sounded like a L1i-trace cache. The fetch-block-optimizer would go to certain hot L1i entries and "optimize" them, with agressive instruction fusion including fusion of non-adjacent instructions. * NextSilicon: Idk any details yet, but they said they handled RVC without increasing latency and that they've found a good solution for implement RVV and especially LMUL, which is a challange in out-of-order designs. * OpenXiangShan: The fastes open-source CPU, is working on doing 2-ahead instruction fetch (the thing Zen5 added). Now that being said, Ventana was bought by Qualcomm, we know the RISC-V team is still alive, but who knows if we'll ever see anything from that outside of Qualcomm? The Tenstorrent Ascalon devboard is way behind schedule and on 12nm TSMC instead of a 4nm node the processor was designed for and is now supposed to clock at 1.38GHz. Though I think the delay has more to do with TT management problems then with the actual design. While the scalar part of OpenXiangShan looks really good, the RVV imolementation is currently basically unusable. They want to have fix for the problems until the end of the year, but we'll have to see.
- philipportner 1mo agoHow do you keep up with such information? Any sources you could recommend? Closest I know would be SemiAnalysis
- rjzzleep 1mo ago