4 ms·
I'm not entirely sure that intel is in the best position to make consumer grade specialty "AI" PC because they make CPUs not GPUs. Maybe that's a bad interpreta
by sovietmudkipz 3y ago
I'm not entirely sure that intel is in the best position to make consumer grade specialty "AI" PC because they make CPUs not GPUs. Maybe that's a bad interpretation on my end since intel has afforded graphics for computers without a dedicated GPU.
Side note: as a hobbyist game dev I've learned about shaders. A shader is simply a program written for a GPU. In game dev we talk about shaders for graphics output but one can write generic shaders too (look up GPGPU). My eyes were opened to the concept that there are programs written for CPU and programs written for GPU.
A CPU typically has <=12 cores and each core computes very quickly. A GPU typically has <=1024 cores and each core computes quickly (but not as quick as a CPU core). It follows that a program ideal for a GPU is one that benefits massive parallelization. I'm still noodling on what that ideal program could be, outside of the obvious graphics processing.
Sharing in case others hadn't realized this tidbit and/or someone more experienced can share their thoughts.
- klohto 3y agoI got news for you brother, maybe one day will use it somewhere, matrix computations
- gumballindie 3y ago> I'm still noodling on what that ideal program could be, outside of the obvious graphics processing. Well machine learning does its computations on GPUs much like shaders. It's basically the same thing with a different function in the kernel: https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.... "CUDA C++ extends C++ by allowing the programmer to define C++ functions, called kernels, that, when called, are executed N times in parallel by N different CUDA threads, as opposed to only once like regular C++ functions."
- derstander 3y ago> I'm not entirely sure that intel is in the best position to make consumer grade specialty "AI" PC because they make CPUs not GPUs. Intel makes (albeit recently) the Arc series of dedicated GPUs. I’m not a graphics programmer so it’s not quite clear how the logic units compare between nVidia, AMD, and Intel, but this link suggests thousands of shading units/cores. https://wccftech.com/intel-arc-a770-a750-a580-graphics-cards-official-specs-unveiled-up-to-32-xe-cores-16-gb-gddr6-2-1-ghz-clocks/amp/ https://wccftech.com/intel-arc-a770-a750-a580-graphics-cards...
- tmccrary55 3y agoMy experience with them is Arc is cheap and low end (but better than Intels previous stuff). Kind of like the new S3 Virge.
- derstander 3y agoThat definitely meshes with what I’ve seen with respect to gaming performance. Any experience with Arc for GP-GPU using e.g. Intel’s DPC++? Though certainly nVidia’s CUDA is by and far the major player in that domain. https://www.intel.com/content/www/us/en/docs/dpcpp-cpp-compiler/developer-guide-reference/2023-2/overview.html https://www.intel.com/content/www/us/en/docs/dpcpp-cpp-compi...
- brucethemoose2 3y ago> they make CPUs not GPUs They do make GPUs! They've made some beefy iGPUs in the past, like Broadwell with eDRAM, and they make discrete consumer GPUs like the very reasonably priced 16GB Arc A770 now, and they make 128GB server GPUs! They have some very large GPUs, like Falcon Shores and the Battlemage enthusiast die, on their roadmap. Also, AI is not always as "core happy" as you would think. For instance, llama.cpp can essentially saturate a DDR5 RAM bus generating tokens on a relatively modest CPU, hence a huge IGP would bring limited benefits in that specific phase. Its also a relatively "serial" operation since the next tokens depend on the previous once, hence its hard to get good GPU utilization without serving multiple clients in parallel. And other "AI chip" designs have diverged from GPUs, like this one: https://fuse.wikichip.org/news/3256/centaur-new-x86-server-processor-packs-an-ai-punch/ https://fuse.wikichip.org/news/3256/centaur-new-x86-server-p... > The NCORE itself has a very interesting design. It’s a whopping 32,768-bits wide SIMD VLIW machine. Its seemingly one giant core!
- Brusco_RF 3y agoWow, you sent me down a serious rabbit hole. That core looks awesome, and the ring bus blew my mind [1]. What is the practical difference between a very wide SIMD processor vs very many single-instruction processors? Is there any? They say this processor is comprised of 16 "slices", each 265B SIMD + 1 MiB cache. Is that different at all from having 16 265B processors? [1] https://en.wikichip.org/wiki/centaur/microarchitectures/cha#Ring https://en.wikichip.org/wiki/centaur/microarchitectures/cha#...
- Dylan16807 3y agoThey all run the same instruction, so I'm not actually sure if the slices matter much.
- brucethemoose2 3y agoA huge SIMD core should have a die area advantage (which in turn gives it other advantages) over a bunch of little processors of the same "width" executing instructions independently. And there is less overhead. In exchange, it can't process different instruction streams simultaneously. It has to perform the same operations over a giant chunk data, or otherwise "waste" the huge SIMD width. Another highlight of this thing: > There is also extensive support for predication with 8 predication registers. The unit is optimized for 8-bit integers (9-bit calculations) From everything I read, NCore would have been a low price LLM monster. Centaur would probably be alive and selling them like hotcakes if they came out with it now, instead if then.
- Arelius 3y ago> A CPU typically has <=12 cores and each core computes very quickly. A GPU typically has <=1024 cores and each core computes quickly (but not as quick as a CPU core). That's not entirely accurate... Cuda "cores" are mostly analogous to a SIMD lanes, which iirc on Nvidia are about 32 registers wide, so 1024/32 = 32 actual individual compute units... Similarly, most Intel CPUs support AVX which can be 16, or even 32 lanes wide in the case of AVX-512. Which gives you 384 lanes. There are of course many other details that make the comparison difficult... Of course the essence of your thinking is in the right direction, but it's just that, in some ways, CPUs and GPUs are have converged more then may be obvious at fitst glance. I will however take issue with the point that Intel may not be well placed, due to the similarities, along with Intel's experience with GPUs, and projects like Larrabee which eventually became Xeon Phi, they certainly have the technical expertise. Now, if they are institutionally well suited for the task, we'll just have to wait and find out.
- rthomas6 3y ago> I'm still noodling on what that ideal program could be, outside of the obvious graphics processing. Basically any kind of big linear algebra workload. Things where you have a big vector/matrix and you need to do some operations to the whole thing. This applies to graphics processing, but also to all DSP in general: RF processing, image manipulation, video encoding, etc. Audio stuff tends to not need it because you only need to process 44000 samples per second, and CPUs are fast enough for that. AI stuff, I understand, uses a lot of linear algebra, so GPUs are good for that workload also.
- jdietrich 3y agoIntel do make GPUs. They're relatively new to the high-performance space, but they have been making iGPUs for a very long time. Their success in the dGPU market is far from certain, but they're making a credible effort. More to the point, we use GPUs because that's what we have, not because they're optimal. Intel have a lot of experience in designing special-purpose accelerators, with perhaps the best example being Quick Sync - a low-end Intel CPU outperforms pretty much anything at video encoding and decoding. Inferencing is a very different workload to training and there is substantial potential for on-CPU accelerators to massively improve the performance and efficiency of inferencing tasks.