10 ms·
The funny thing to me is that so much of the "AI software ecosystem" is just PyTorch. You don't need to develop some new framework and make it popular. You don'
by lacker 3y ago
The funny thing to me is that so much of the "AI software ecosystem" is just PyTorch. You don't need to develop some new framework and make it popular. You don't need to support a zillion end libraries. Just literally support PyTorch.
If PyTorch worked fine on Intel GPUs, a lot of people would be happy to switch.
- latchkey 3y agoThis is a big reason why AMD did this deal with PyTorch... https://pytorch.org/blog/experience-power-pytorch-2.0/ https://pytorch.org/blog/experience-power-pytorch-2.0/
- shihab 3y agoBut you can't support Pytorch without a proper foundation in place. They don't need to support zillion _end_ libraries, sure, but they do need to have at least a very good set of standard libraries, equivalent of Cublas, Curand etc. And they don't. My work recently had me working with rocRAND (Rocm's answer to Curand). It was frankly pretty bad- the design, performance (50% slower in places that don't make any sense because generating random numbers is not exact that complicated), and documentation (God it was awful). Now, that's a small slice of the larger pie. But imagine if this trend continues for other libraries.
- lumost 3y agoFolks also underestimate how complex these libraries are. There are dozens of projects to make BLAS alternatives which give up after ~3-6 months when they realize that this project will take years to be successful.
- eternityforest 3y agoHow does that work? Why not pick up where the previous team left off instead of everyone starting new ones? Or are they all targeting different backends and hardware?
- haltist 3y agoIt's a compiler problem and there is no money in compilers [1]. If someone made an intermediate representation for AI graphs and then wrote a compiler from that intermediate format into whatever backend was the deployment target then they might be able to charge money for support and bug fixes but that would be it. It's not the kind of business anyone wants to be in so there is no good intermediate format and compiler that is platform agnostic. 1: https://tinygrad.org/ https://tinygrad.org/
- whatshisface 3y agoJAX is a compiler.
- deleted 3y ago[deleted]
- haltist 3y agoSo are TensorFlow and PyTorch. All AI/ML frameworks have to translate high-level tensor programs into executable artifacts for the given hardware and they're all given away for free because there is no way to make money with them. It's all open source and free. So the big tech companies subsidize the compilers because they want hardware to be the moat. It's why the running joke is that I need $80B to build AGI. The software is cheap/free, the hardware costs money.
- singhrac 3y agoGenerating random numbers is a bit complicated! I wrote some of the samplers in Pytorch (probably replaced by now) and some of the underlying pseudo-random algorithms that work correctly in parallel are not exactly easy... running the same PRNG with the same seed on all your cores will produce the same result, which is probably NOT what you want from your API. But, to be honest, it's not that hard either. I'm surprised their API is 2x slower, Philox is 10 years old now and I don't think there's a licensing fee?
- eternityforest 3y agoI wonder if the next generation chips are going to just have a dedicated hardware RNG per-core if that's an issue?
- lelanthran 3y agoWhy bother? It's not the generation that matters so much, it's the gathering of entropy, which comes from peripherals and not possible to generate on-die. If you don't need cryptographically secure randomness, you still want the entropy for generating the seeds per thread/die/chip.
- eternityforest 3y agoIt absolutely is possible to generate entropy on-die, assuming you actually want entropy and not just a unique value that gets XORed with the seed, so you can still have repeatable seeds. Pretty much every chip has an RNG which can be as simple as just a single free running oscillator you sample
- lelanthran 3y ago> Pretty much every chip has an RNG which can be as simple as just a single free running oscillator you sample Every chip may have some sort of noise to sample, but they are nowhere near good sources of entropy. Entropy is not a binary thing (you either have it or don't), it's a spectrum and entropy gathered on-die is poor entropy. Look, I concede that my knowledge on this subject is a bit dated, but the last time I checked there were no good sources of entropy on-die for any chip in wide use. All cryptographically secure RNGs depend on a peripheral to grab noise from the environment to mix into the entropy pool. A free-running oscillator is a very poor source of entropy.
- nradov 3y agoInstead of generating pseudorandom numbers you can just download files of them. https://archive.random.org/ https://archive.random.org/
- hcrean 3y agoOr you could just re-use the same number; no one can prove it is not random. https://xkcd.com/221/ https://xkcd.com/221/
- slavik81 3y agoIf you haven't already, please consider filing issues on the rocrand GitHub repo for the problems you encountered. The rocrand library is being actively developed and your feedback would be valuable for guiding improvements.
- shihab 3y agoAppreciate it, will do.
- jart 3y agoI honestly don't see why it's so hard. On my project we wrote our own gemm kernels from scratch so llama.cpp didn't need to depend on cublas anymore. Only took a few days and a few hundred lines of code. We had to trade away 5% performance.
- soulbadguy 3y agoFor a given set kernels, and a limited set of architectures, the problem is relatively easy. But covering all the important kernels acros all the crazy architecture out there and with relatively good performance and numerical accuracy ... Much harder
- rightbyte 3y ago> generating random numbers You can't bench implementations of random numbers against each other purely on execution speed. A better algorithm (better statistical properties) will be slower.
- codetrotter 3y agoI have the fastest random number generator in the world. And it works in parallel too! https://i.stack.imgur.com/gFZCK.jpg https://i.stack.imgur.com/gFZCK.jpg
- shihab 3y agoYeah. In this instance, I was talking about the same algorithm (phillox), the difference is purely in implementation.
- zozbot234 3y agoPyTorch includes some Vulkan compat already (though mostly tested on Android, not on desktop/server platforms), and they're sort of planning to work on OpenCL 3.0 compat, which would in turn lead to broad-based hardware support via Mesa's RustiCL driver. (They don't advertise this as "support" because they have higher standards for what that term means. PyTorch includes a zillion different "operators" and some of them might be unimplemented still. Besides performance is still lacking compared to CUDA, Rocm or Metal on leading hardware - so only useful for toy models.)
- singhrac 3y agoJust to point out it does, kind of: https://github.com/intel/intel-extension-for-pytorch https://github.com/intel/intel-extension-for-pytorch I've asked before if they'll merge it back into PyTorch main and include it in the CI, not sure if they've done that yet. In this case I think the biggest bottleneck is just that they don't have a fast enough card that can compete with having a 3090 or an A100. And Gaudi is stuck on a different software platform which doesn't seem as flexible as an A100.
- spacemanspiff01 3y agoThey could compete on ram, if the software was there. Just having a low cost alternative to the 4060ti would allow them to break into the student/hobbies/open source market. I tried the a770, but returned it. Half the stuff does not work. They have the CPU side and GPU development on different branches (GPU seems to be ~6 months behind CPU) and often you have to compile it yourself, (if you want torchvision or torchaudio) it also currently on 2.0.1 of pytorch so somewhat lagging, and does not have most of the performance analysis software available. You also, do need to modify your pytorch code, often more than just replacing cuda for xpu as the device. They are also doing all development internally, then pushing intermittently to public. A lot of this would not be as bad if there was a better idea of feature timeline, or if they made their CI public. (Trying to build it myself involved a extremely hacky bash script, that inevitably failed halfway through.)
- selfhoster11 3y agoThe amount of VRAM is the absolute killer USP for the current large AI model hobbyist segment. Something that had just as much VRAM as a 3090 but at half the speed and half the price would sell like hot cakes.
- eyegor 3y agoYou are describing the ebay market for used nvidia tesla cards. The k80, p40, or m40 are widely available and sell for ~$100 with 24gb vram. The m10 even has 32gb! The problem for ai hobbyists is it won't take long to realize how many apis use the "optical flow" pathways and so on nvidia they'll only run at acceptable speeds on rtx hardware, assuming they run at all. Cuda versions are pinned to hardware to some extent.
- ndneighbor 3y agoOneAPI isn't bad for PyTorch, the performance isn't there yet but you can tell it's an extremely top priority for Intel.
- flakiness 3y agoIntel has to do it by themselves. NVIDIA just lets Meta/OpenAI/Google engineers do it for them. Such a handicapped fight.
- joe_the_user 3y agoThat's because CUDA is a clear, well-functioning library and Intel has no equivalent. It makes any "you just have to get Pytorch working" a little less plausible.
- seanhunter 3y agoIt wasn’t always like this. Nvidia did the initial heavy lifting to get cuda off the ground to a point where other people could use it.
- seanhunter 3y agoBut this is the the thing. Speaking as someone who dabbles in this area rather than any kind of expert, it’s baffling to me that people like Intel are making press releases and public statements rather than (I don’t know) putting in the frikkin work to make performance of the one library that people actually use decent. You have a massive organization full of gazillions of engineers many of whom are really excellent. Before you open your mouth in public and say something is a priority, deploy a lot of them against this and manifest that priority by actually doing the thing that is necessary so people can use your stuff. It’s really hard to take them seriously when they haven’t (yet) done that.
- pas 3y agoYou know how it works. The same busybodies who are putting out this useless noise releases are the ones who squandered Intel's lead, and now are patting themselves on the back for figuring out that with this they'll again be on top for sure! There was a post on HN a few months ago about how Nvidia's CEO still has meetings with engineers in the trenches. Contrast that with what we know of Intel, which is not much good, and a lot of bad. (That they are notoriously not-well-paying, because they were riding on their name recognition.)