5 ms·
12GB memory -.- I feel like _anyone_ who can pump out GPU's with 24GB+ of memory that are capable to use for py-stuff would benefit greatly. Even if it's not
by Implicated 2y ago
12GB memory
-.-
I feel like _anyone_ who can pump out GPU's with 24GB+ of memory that are capable to use for py-stuff would benefit greatly.
Even if it's not as performant as the NVIDIA options - just to be able to get the models to run, at whatever speed.
They would fly off the shelves.
- cowmix 2y ago100% - this could be Intel's ticket to capture the hearts of developers and then everything else that flows downstream. They have nothing to lose here -- just do it Intel!
- evanjrowley 2y agoMaybe that's not too bad for someone who wants to use pre-existing models. Their AI Playground examples require at minimum an Intel Core Ultra H CPU, which is quite low-powered compared to even these dedicated GPUs: https://github.com/intel/AI-Playground https://github.com/intel/AI-Playground
- elorant 2y agoWould it though? How many people are running inference at home? Outside of enthusiasts I don't know anyone. Even companies don't self-host models and prefer to use APIs. Not that I wouldn't like a consumer GPU with tons of VRAM, but I think that the market for it is quite small for companies to invest building it. If you bother to look at Steam's hardware stats you'll notice that only a small percentage is using high-end cards.
- ModernMech 2y agoIt's a chicken and egg scenario. The main problem with running inference at home is the lack of hardware. If the hardware was there more people would do it. And it's not a problem if "enthusiasts" are the only ones using it because that's to be expected at this stage of the tech cycle. If the market is small just charge more, the enthusiasts will pay it. Once more enthusiasts are running inference at home, then the late adopters will eventually come along.
- m00x 2y agoMac minis are great for this. They're cheap-ish and they can run quite large models at a decent speed if you run it with an MLX backend.
- alganet 2y agomini _Pro_ are great for this, ones with large RAM upgrades. If you get the base 16GB mini, it will have more or less the same VRAM but way worse performance than an Arc. If you already have a PC, it makes sense to go for the cheapest 12GB card instead of a base mac mini.
- tokioyoyo 2y agoThis is the weird part, I saw the same comments in other threads. People keep saying how everyone yearns for local LLMs… but other than hardcore enthusiasts it just sounds like a bad investment? Like it’s a smaller market than gaming GPUs. And by the time anyone runs them locally, you’ll have bigger/better models and GPUs coming out, so you won’t even be able to make use of them. Maybe the whole “indoctrinate users to be a part of Intel ecosystem, so when they go work for big companies they would vouch for it” would have merit… if others weren’t innovating and making their products better (like NVIDIA).
- m00x 2y agoYou can just use a CPU in that case, no? You can run most ML inference on vectorized operations on modern CPUs at a fraction of the price.
- marcyb5st 2y agoMy 7800x says not really. Compared to my 3070 it feels so incredibly slow that gets in the way of productivity. Specifically, waiting ~2 seconds vs ~20 for a code snippet is much more detrimental to my productivity than the time difference would suggest. In ~2 seconds I don't get distracted, in ~20 seconds my mind starts wandering and then I have to spend time refocusing. Make a GPU that is 50% slower than a 2 generations older mid-range GPU (in tokens/s) but on bigger models and I would gladly shell out 1000+$. So much so that I am considering getting a 5090 if nVdia actually fixes the connector mess they made with 4090s or even a used v100.
- refulgentis 2y agoI don't understand, make it slower so it's faster?
- magicalhippo 2y agoMy 2080Ti at half speed would still beat the crap out of my 5900X CPU for inference, as long as the model fits in VRAM. I think that's what GP was alluding to.
- m00x 2y agoI'm running codeseeker 13B model on my macbook with no perf issues and I get a response within a few seconds. Running a specialist model makes more sense on small devices.
- bongodongobob 2y agoI don't know a single person in real life that has any desire to run local LLMs. Even amongst my colleagues and tech friends, not very many use LLMs period. It's still very niche outside AI enthusiasts. GPT is better than anything I can run locally anyway. It's not as popular as you think it is.
- dimensi0nal 2y agoThe only consumer demand for local AI models is for generating pornography
- treprinum 2y agoHow about running your intelligent home with a voice assistant on your own computer? In privacy-oriented countries (Germany) that would be massive.
- magicalhippo 2y agoThis is what I'm fiddling with. My 2080Ti is not quite enough to make it viable. I find the small models fail too often, so need larger Whisper and LLM models. Like the 4060 Ti would have been a nice fit if it hadn't been for the narrow memory bus, which makes it slower than my 2080 Ti for LLM inference. A more expensive card has the downside of not being cheap enough to justify idling in my server, and my gaming card is at times busy gaming.
- serf 2y agoabsolutely wrong -- if you're not clever enough to think of any other reason to run an LLM locally then don't condemn the rest of the world to "well they're just using it for porno!"
- knowitnone 2y agoso you're saying that a huge market?!
- throwaway48476 2y ago
- rafaelmn 2y agoYou can get that on mac mini and it will probably cost you less than equivalent PC setup. Should also perform better than low end Intel GPU and be better supported. Will use less power as well.