5 ms·
Then this is SD for Apple silicon users. 13B runs on my m1 air at 200-300ms/token using llama.cpp. Outputs feel like original GPT-3, unlike any of the competito
by bestcoder69 4y ago
Then this is SD for Apple silicon users. 13B runs on my m1 air at 200-300ms/token using llama.cpp. Outputs feel like original GPT-3, unlike any of the competitors I’ve tried. Granted- non-scientific first impressions.
- j45 4y agoAgreed. For those who have been quietly sitting with a base Mac Studio, or a reasonably capable Mac Mini. The possibilities changed on some fronts, but GPT's extremely low price on their API remains a good option.
- aaomidi 4y agoDifference is chatgpt is not privacy friendly.
- staticautomatic 4y agoIs it still not privacy friendly on Azure?
- aaomidi 4y agoAzure has access to your queries. Running locally really is the only way of having a privacy friendly LLM.
- fastball 4y agoI wonder if there is some way you could do an E2EE LLM SaSS? Guess more work probably needs to be done into homomorphic encryption for that.
- jamiek88 4y agoThis would actually be a great use case for homomorphic encryption. I’m behind on what the state of the art is in that though, but my mind immediately went there.
- aaomidi 4y agoI don’t think homomorphic encryption works like that
- jamiek88 4y agoWell that was helpful.
- hobs 4y agoIt's insanely slow and probably wont be fast for anything for a long time.
- speedgoose 4y agoFrom my engineering point of view, the state of the art of homomorphic encryption can do some maths very very slowly at a huge cost and can’t be used yet for any real use case. A very cool research topic, but it’s much simpler to run our software locally if you don’t want to leak your data.
- fullsend 4y agoIt feels like you don’t need encryption here, you just need a business model that doesn’t keep user accounts and sell the query data.
- deleted 4y ago[deleted]
- j45 4y agoIs the chatgpt api paid to openai not private to openai? I understand azure has its own chatgpt api.
- nordsieck 4y ago> or a reasonably capable Mac Mini IMO, Apple's habit of cheaping out on ram is maddening. Although perhaps this is a sign for me to pick up a m2 pro mini with extra ram.
- j45 4y agoAgreed. Apple might have sped up the ram addressing by building it in this time. The Base mac studio with 32 gigs of ram can be a good value from a ram perspective. Let’s see if the Mac studio is discontinued or updated.
- nordsieck 4y ago> Let’s see if the Mac studio is discontinued or updated. I do hope that it's updated. I could see them reserving the top end chip for the Mac Pro, though. Or maybe make it a dual socket? In any case, I much prefer the ability to run an Apple computer with a proper cooling setup. I know that cooling pads for laptops exist, but ultimately it's a bit of a janky solution compared to actually adequately cooling the system properly.
- j45 4y agoYou’re bang on with this comment. The Mac studio is surprisingly quiet and well cooled. I haven’t had to install a fan control utility as of yet. I thought I’d be trying one of these out and sell it if my setup didn’t work, but it’s the first desktop I’m starting to consider especially if I want to leave a workload running on it instead of a decade of carrying a lot of horsepower around with me. Instant on for 3 monitors, no need for docking station or USB hubs, everything is plugged in and works, I can run a virtual cam obs setup easily. The studio is a great upgrade on Mac mini, which might cannibalize the Mac pro sales and why it might not get toasted or refreshed. I’ll F the max studio was a way to put a dent in hackkntoshes, it makes sense. The M2 mini looks super decent but quickly goes up in price when ram and ssd is setup to match a studio. For my purposes the ultra wasn’t that much faster than the M1 max on because the software has to be optimized to benefit from it. Maybe some of this ML stuff will be shortly.
- davidy123 4y agoYou can put together a 192GB 16+ core x86 system with 16GB CUDA card and multi TB of fast storage for $2000ish. I'm not trolling, just wondering why people go to "Apple" all the time, when other approaches may be better for this kind of work. Yes, the CPU - RAM interface is not as fast, but the CUDA card is much faster, and the large cheap memory makes some things a lot more practical. If I'm not mistaken, in these approaches, GPU + VRAM + CPU + RAM are used in conjunction, so it all adds up to quite a bit more powerful system for the same amount of money if working with them is a main goal. If Apple had expandable RAM it would be a different story.
- smoldesu 4y agoScrew that - the other day I realized that it's cheaper to buy an Intel A770 ($350) with 16gb of memory than it is to upgrade a Mac Mini with 16gb ($400) of extra memory. Apple's optimization here is nice for the people who own their hardware, but it's totally silly to read through the comments promising the end of CUDA.
- simonw 4y agoIn my case I'm interested in AI, but not quite enough to spend money on a whole separate computer for exploring ML work. I want a really great single laptop I can use for everything.
- j45 4y agoI’m as much PC/Linux as I am Mac. Lots of experience putting together more computers than I can count, or having them put together. I think there might be a different use case for me but to compare with your economics - I picked up a base mac studio (m1 max with 32 gb) for about $1370 USD with 8 months warranty remaining. It’s letting me test my daily setup to see if everything can run on Apple silicon yet without intervention.. if not, I’ll be able to get rid of it at little to no loss, and decide if I want to carry that much computing power in a laptop or not, or head on to other options. My interest currently is computational power, per watt. Not much comes close to Apple Silicon. The cost of a loaded pc can quickly outstrip other options with electricity costs included so it needs to be useful in any case. The integrated speed of the Apple silicon, ram, and ssd is a little astonishing. More than I expected to admit. I don’t know if there’s anything like it on PC. If Apple silicon supported eGPUs it would interest me. Comparable PCs have described are generally power inefficient. Still, the system you’re laying out is interesting, especially the ram, mind sharing a bill of materials?
- deleted 4y ago[deleted]
- deleted 4y ago[deleted]
- ddren 4y agoThey have recently merged support for x86. I get 230ms/token on the 13B model on a 8 core 9900k under WSL2.
- qumpis 4y agoWhat's your ram usage for this?
- ddren 4y agoThe (quantized) 13B model is 7.6 GB on disk and the program uses around 8 GB to run. It runs without hitting the swap with just 9 GB assigned to WSL2.
- JoeMattie 4y agoI've got the (4-bit quantized) 65B param model running at somewhat acceptable speed on an i9-7900. It uses around 55GB of RAM.
- MacsHeadroom 4y ago13B in 4bit runs on your average 8GB VRAM consumer video card at fast enough speeds. The GP's info is outdated. Apparently they don't realize how quickly LLaMA is being developed on. It's just like the early days of SD.
- smy20011 4y agoMaybe out-dated by 1 day LOL.
- zamadatix 4y agollama.cpp runs on the CPU anyways, the whole conversation is on a false premise. The only hard thing about running it on consumer hardware is the quantization step (where the largest RAM requirements come in) and that's not even a technical problem since you could just redistribute the quantized version if it weren't for the license of the model.
- tomp 4y agoSame here. running 4-bit LLaMa 13B on my MacBook Air M2: "possible resolutions to the Fermi paradox, the question of where is the evidence for extra-terrestrial civilizations: - " answer: possible resolutions to the Fermi paradox, the question of where is the evidence for extra-terrestrial civilizations: - 1. We are looking in the wrong direction and in the wrong frequency bands for signals from intelligent life in the galaxy. - 2. We are looking in the right frequency bands, but the technological level of intelligent life in the galaxy is not high enough to be a threat to us. - 3. We are looking in the right direction and in the right frequency bands, but the probablity of interstellar travel is so small that we are not likely to meet civilizations. - 4. We are looking in the right direction and in the right frequency bands, and the probability of interstellar travel is not small, but we are unlikely to meet civilizations at our level of technological development. - 5. We are looking in the right direction and in the right frequency bands, and the probability of interstellar travel is not small, and we are likely to meet civilizations at our level of technological development. - 6. We are looking in the right direction and in the right frequency bands, and the probability of interstellar travel is not small, and we are likely to meet civilizations at our level of technological development. I would be inclined to think that option 5 or 6 are the most likely. Given the nature of the universe, I think that we are more likely to encounter civilizations in the earliest stages of their development, than those that are further developed. I would think that the big question of the next 10 years is whether SETI will be able to detect an ET signal. I think that the chances of being able to do so are less than even. The reason is that a lot of the technical issues that SETI has been addressing are becoming less of a problem, because the Earth-based radio telescopes are being joined by many smaller radio telescopes that are being sent into space. So, from a technical standpoint, the ability to look for ET signals is becoming easier and easier. The problem is that any signals that are sent to Earth that come from technologically advanced civilizations are likely to be only a few percent of the power of the ET signals that are coming from sources at our level of technological development. So, for example, SETI@home is looking for