10 ms·
Promote and proliferate local LLMs. If you use GPT, you're giving OpenAI money to lobby the government so they'll have no competitors, ultimately screwing your
by PostOnce 3y ago
Promote and proliferate local LLMs.
If you use GPT, you're giving OpenAI money to lobby the government so they'll have no competitors, ultimately screwing yourself, your wallet, and the rest of us too.
OpenAI has no moat, unless you give them money to write legislation.
I can currently run some scary smart and fast LLMs on a 5 year old laptop with no GPU. The future is, at least, interesting.
- john2x 3y agoCare to share some links? My lack of GPU is the main blocker for me from playing with local-only options. I have an old laptop with 16GB RAM and no GPU. Can I run these models?
- PostOnce 3y agohttps://github.com/ggerganov/llama.cpp https://github.com/ggerganov/llama.cpp https://huggingface.co/TheBloke https://huggingface.co/TheBloke There's a LocalLLaMA subreddit, irc channels, and a whole big community around the web working on it on GitHub nd elsewhere. edit: I forgot to directly answer you: yes you can run these models. 16GB of plenty. Different quantizations give you different amounts of smarts and speed. There are tables that tell you how much RAM is needed per which quantization you choose, as well as how fast it can produce results (ms per token). e.g. https://github.com/ggerganov/llama.cpp#quantization https://github.com/ggerganov/llama.cpp#quantization where RAM required a little more than the file size, but there are tables that list it explicitly which I don't have immediately at hand.
- tensor 3y agoA reminder that llama isn't legal for the vast majority of use cases. Unless you signed their contract and then you can use it only for research purposes.
- rvcdbn 3y agoWe don’t actually know that it’s not legal. The copyrightability of model weights is an open legal question right now afaik.
- tensor 3y agoIt doesn't have to be copyrightable to be intellectual property.
- actionfromafar 3y agoPatents? Trademark? What do you mean?
- Rexxar 3y agoMaybe this: https://en.wikipedia.org/wiki/Database_right https://en.wikipedia.org/wiki/Database_right but it doesn't exist in every countries.
- twbarr 3y agoNo, but what is it? Not your lawyer, not legal advice, but it's not a trade secret, they've given it to researchers. It's not a trademark because it's not an origin identifier. The structure might be patentable, but the weights won't be. It's certainly not a mask work. It might have been a contract violation for the guy who redistributed it, but I'm not a party to that contract.
- Art9681 3y agoI'm going to play devil's advocate and state that a lot of what you mentioned will be relevant to a tiny part of the world that has the means to enforce this. The law will be forced to change as a response to AI. Many debates will be had. Many crap laws will be made by people grasping at straws but it's too late. Putting red tape around this technology puts that nation at a technological disadvantage. I would go as far as labeling a national security threat. I'm calling it now. Based on what I see today. Europe will position itself as a leader in AI legislation, and its economy will give way to the nations that want to enter the race and grab a chunk of the new economy. It's a Catch 22. You either gimp your own technological progress, or start a war with a nation that does not. Pretty sure Russia and China don't really care about the ethics behind it. There are plenty of nations capable enough in the same boat. Now what? OK, so in some hypothetical future China has an uncensored model with free reign over the internet. The US and Europe has banned this. What's stopping anyone from running the Chinese model? There isn't enough money in the world to enforce software laws. How long have they tried to take down The Pirate Bay? Pretty much every permutation of every software that's ever been banned can be found and ran with impunity if you have the technical knowledge to do so. No law exists that can prevent that. If it did, OpenAI wouldn't exist.
- PostOnce 3y agoOpenLLaMA is though. https://github.com/openlm-research/open_llama https://github.com/openlm-research/open_llama All of these are surmountable problems. We can beat OpenAI. We can drain their moat.
- donw 3y agoFor the above, are the RAM figures system RAM or GPU?
- PostOnce 3y agoCPU RAM
- kordlessagain 3y ago> We can drain their moat. I've got an AI powered sump pump if you need it.
- ignoramous 3y agoThey most certainly don't need / deserve the snark, to be sure, on hacker news of all places.
- tensor 3y agoAbsolutely, 100% agree. I just wouldn't touch the original LLaMA weights. There are many amazing open source models being built that should be used instead.
- niemandhier 3y agoIt’s not clear if their license terms would hold, for the moment just act and worry later. Update: That is only true for the legal system I am currently residing in. No idea about e.g. the US.
- lhl 3y agoThis is the most well-maintained list of commercially usable open LLMs: https://github.com/eugeneyan/open-llms https://github.com/eugeneyan/open-llms MPT, OpenLLaMA, and Falcon are probably the most generally useful. For code, Replit Code (specifically replit-code-instruct-glaive) and StarCoder (WizardCoder-15B) are the current top open models and both can be used commercially.
- jstummbillig 3y agoJust a heads up: If you are more interested in being effective than being an evangelist, beware. While you can run all kinds of GPTs locally, GPT-4 still smokes everything right now – and even it is not actually good enough to not be a lynchpin for a lot of cases yet.
- slaymaker1907 3y agoI guess ignoring copyright and treating the whole internet as your training data does have its advantages.
- dcow 3y agoYes? That’s the point. Who cares about an outdated concept that has no digital analog? All the artists have moved on already #midjourney.
- bombolo 3y agoWhen mirosoft will open up all of their source code, I will agree with you.
- raxxorraxor 3y agoNo, I doubt artists have moved on. And if they want no artificial gatekeeper, than it is #stablediffusion instead of #midjourney. I would argue that it creates better images too.
- logicchains 3y ago>GPT-4 still smokes everything right now Not if you want it to write adult (graphically pornographic or violent) content.
- tudorw 3y agohttps://gpt4all.io/index.html https://gpt4all.io/index.html
- yard2010 3y agoKeep in mind it doesn't relate to GPT4, the 4 in the name is for, not four. But I should try it. TBH openAI shady practices and MS behind them is just an anti trust waiting to happen and I don't want a part in this dystopia
- moffkalast 3y ago16GB of RAM can fit a 5 bit 13B model at best, they're second dumbest class of LLama model. If Open Orca turns out any good than that might be enough for the time being, but you'll need more RAM to use anything serious. Here's a handy model comparison chart (this is a coding benchmark, so coding-only models tend to rank higher): https://i.imgur.com/AqSjjj2.jpeg https://i.imgur.com/AqSjjj2.jpeg
- PostOnce 3y agoYour benchmark lacks the current #2 https://github.com/nlpxucan/WizardLM/tree/main/WizardCoder https://github.com/nlpxucan/WizardLM/tree/main/WizardCoder It beats Claude and Bard. You could probably get a 4bit 15B model going in 16GB of RAM and be approaching GPT4 in capability. ...on an old laptop, lol Let's eat OpenAI's lunch! They deserve it for trying to steal this tech by "privatizing" a charity, hiding scientific data that was supposed to be shared with us by said charity whose purpose was to help us all, and dishonestly trying to persuade the government not to let us compete with them.
- moffkalast 3y agoYeah I mean I wouldn't really include coding models in this list since they're not general purpose models and have an obvious fine tuning edge compared to the rest. But WizardCoder is definitely something to look at as a Copilot replacement. I'd post a more well rounded benchmark but the problem is that all non-coding benchmarks are currently more or less complete garbage, especially the Vicuna benchmark that rates everything as 99.7% GPT 3.5 lol.
- PostOnce 3y agoThe benchmark you linked was to "programming performance", not generic LLM "intelligence". The situation for the little guy is wildly better than most people imagine.
- moffkalast 3y agoYep, that's what I'm saying, programming performance is seemingly very indicative of model inteligence (assuming it's tuned well enough to be able to run the benchmark at all). Coding is an exercise in problem solving and abstract thinking after all. There are exceptions of course, as there are a few models (e.g. Vicuna, Baize) that don't do well at coding at all but otherwise perform well for chat, and the coding models I mentioned that game the benchmark by sacrificing performance in all other areas. If you exclude those, it's very a accurate overall reasoning level comparison, at least it fits most to what I've seen their performance was for various tasks when testing out individual models. The only other valid benchmark that isn't coding are the SAT and LSAT tests that OpenAI runs on all of their models, but afaik there isn't an open version that would be widely used.
- gowld 3y agoThere's no need to run locally if you aren't utilizing 8 hrs/day. You can rent time on a hosted GPU, sharing a hosted model with others.
- joeythedolphin 3y agoGreat point -- I was thinking of renewing my $20/subscription but I will keep it cancelled. We must not fund AI propaganda machines.
- ed_mercer 3y agoForgive me as I’m out of the loop. What propaganda are you referring to?
- NortySpock 3y agoThese ChatGPT tools allow anyone to write short marketing and propaganda prompts. They can then take the resulting paragraphs of puffery and post them using bots or sock puppets to whatever target community to create the illusion of action, consensus, conflict, discussion or dissention. It used to be this took a few people to come up with writing actual responses to forum posts all day, or marketing operations plans, or pro- or anti-thing propaganda plans. But now, you could astroturf a movement with a GPU, a ChatGPT clone, some bots and vpns hosted from a single computer, a cron job, and one human running it. If you thought disinformation was bad 2 years ago, get ready for fully automated disinformation that can be targeted down to an online community or specific user in an online community...
- kaba0 3y agoI believe a new wave of authentication might come out of this, where it is tied to citizenship for example (or something related to physical reality). Otherwise we will find ourselves in a truly chaotic situation.
- joeythedolphin 3y agoSam tells Congress that AI is so dangerous it will extinct humanity. Why? So Congress can license him and only his buddies. Then get goes to euro and speaks with world leaders to remove consumer protection. Why? So he can mine data without any consequences. He is a narcissistic CEO who lies to win. If you are tired of the past decade of electronic corporate tyranny, abuse, manipulation and lies, then boycott OpenAi (should be named ClosedAi) and support open source, or ethical companies (if there are any).
- dimgl 3y agoI'd love to get into AI and AI development. Where can I start?
- musha68k 3y agoOne has to give them credit for what must be the most grandiose stunt actually landed. And on so many angles! “It just works” - they even got the scientists fully aligned! Fiercely smart industriousness. https://youtu.be/P_ACcQxJIsg?t=5946 https://youtu.be/P_ACcQxJIsg?t=5946
- jonplackett 3y agoNo equity? For real? He really does need an agent if that's the case.
- darkerside 3y agoWow, under penalty of perjury
- whimsicalism 3y agoIf you listen to him talk at any point, you can see him explain why.
- andai 3y agoCan you elaborate on scary smart and fast? It's been a month or two since I've tried but the results were depressingly slow and useless for more or less every task I tried. Every time a model is claimed to be "90% of GPT-3" I get excited and every time it's very disappointing. (On that note, after using GPT-4, GPT-3 now seems disappointing almost every time I interact with it.)
- whimsicalism 3y agoi think the falcon instruct is considered pretty good but if you are expectation set by gpt4 it still will not compare
- PostOnce 3y agoDifferent quantizations can give you a big speedup if you've had "depressingly slow" issues. Even the slowest ones (that fit in RAM) will run at basically interactive speed, not instant, but also not "email speed". I have a laptop with a 2018 CPU and I'm working with them just fine. Text generation style instead of chat style is another avenue that makes the feedback time not so annoying for a developer. at 100ms/token, it's faster than most people type, I think. That's what you might get on an old laptop with a 7B model. There's a useful leaderboard here to help you pick a model: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb... It really depends on your task, lots and lots of natural language type tasks give great results, the models seem to have extensive knowledge of many fields. So for some kinds of Q&A bot (technical or not), for copy blurbs, for fiction, game NPCs, etc, the models (especially 13B and up) can be breathtaking, even moreso considering they run on bottom-dollar consumer hardware (I paid $250 for the laptop I'm developing on). There are of course some things that neither the local LLMs nor GPT4 can do, like create useful OpenSCAD models :) Things keep getting better, newer quantization methods give you more smarts in the same amount of RAM at basically the same speed -- the models are getting better, there are more permissively licensed ones now.
- RossBencina 3y agoHow are you running inference? GPU or CPU? I'm trying to use GPT4All (ggml-based) on 32 cores of E5-v3 hardware and even the 4GB models are depressingly slow as far as I'm concerned (i.e. slower than the GPT4 API, which is barely usable for interactive work). I'd be much obliged if you could point me at a specific quantized model on HF that you think is "fast" and I'll download it and try it out.
- andai 3y agoWhich models are you using and for which tasks? I have found local models largely a waste of time (except for very simple tasks with very heavy prompting). But perhaps there are some recent breakthroughs I haven't seen yet.
- PostOnce 3y agoI'm using a variety of 7 and 13B models (and a 3B one for fast feedback loop debugging) at between 8bit and 4_K_M quantizations. Depending on your pre-prompt, your fine-tune (i.e. which model you downloaded), and your specific task, the results can be startlingly good, it's crazy that you can do this on a $250 laptop. I stay up nights working on it lately, it's so interesting. More importantly, things change by the day. New models, new methods, new software, new interfaces... the possibilities are endless... unless we let OpenAI corrupt our government(s).
- redox99 3y agoI'm surprised you're having such a good time with 7B and 13B models. I find anything below 33B to be almost useless. And only 65B is close to GPT 3.5.
- mdale 3y agoI don't think the "corrupt our government" thing is going to happen . The wave of change is too large the tech is moving too fast and into evey facet of data and software. There is competition globally and locally; a regulatory slow down is unlikely.
- kristianp 3y agoGpt-4 runs on 8 x 220B params[1] and gpt is about 220B params(?). Local LLMs can be good for some tasks, but they are much slower and less capable than the size of model and hardware that openai brings to their apis. Even running a 7B model on the CPU in ggml is much slower than the gpt-3-turbo api, in my experience with a 12th gen i7 intel laptop. [1] GPT4 is 8 x 220B params = 1.7T params: https://news.ycombinator.com/item?id=36413296 https://news.ycombinator.com/item?id=36413296
- Art9681 3y agoIt's been well documented by now that the number of parameters does not necessarily translate to a better model. My guess is that OpenAI has learned a thing or two from the endless papers published daily that your "instance" of the model is not what it seems. They likely have a workflow that picks the best model suitable for your prompt. Some people may get a 13B permutation because it is "good enough" to produce a common answer to a common prompt. Why waste precious compute resources on a prompt that is common? Would it not be feasible to collect the data of the top worldwide prompts and produce a small model that can answer those? Why would OpenAI spend precious compute time on the typical user's "write a short story of...". I would guesstimate that the great majority of prompts are trash. People playing with a toy and amusing themselves. The platform sends those to the trash models. For the other tiny percentage that produces a prompt the size of a paragraph, using the techniques published by OpenAI themselves, they likely get the higher tier models. This is also why I believe many are recently complaining about the quality of the outputs. When your chat history is filled with "have waifu pretend to be my girlfriend" then whatever memory the model is maintaining will be poisoned by the quality of your past prompts. Garbage in, garbage out. I am certain that the #1 priority for OpenAI/Microsoft is lowering the cost of each prompt while satisfying the majority. The majority is not in HN.
- ranguna 3y ago> It's been well documented by now that the number of parameters does not necessarily translate to a better model. That's certainly true, but it's hard to deny the quality of gpt 4. If the issue is the training data, let's just use their training data, it's not like they had to close up shop because of using restricted data. I think the issue is more on the financial side, it must have been extremely expensive to train gpt 4. Open source models don't have that kind of money right now. I'll finance open source models once they are actually good, or show realistic promises of reaching that level of quality on consumer hardware. Until then, open source will open source. I've never bought any kind of subscription or paid api costs to openai, but if gpt 4 finally reached the point where I feel like it's a lot better than just good enough, I'll happily pay for it (while still being on the lookout for open source models that fit my hardware).
- willsmith72 3y agoMy laptop already works too hard doing development and having chrome open, it's just not feasible. A good hosted alternative, sure, but local is not going to scale to the masses.
- PostOnce 3y agoI have a Dell 7490 (intel 8350u cpu) I paid $250 for and I have no trouble running 13B models through a custom interactive interface I wrote as a hobby project in an afternoon. It can still get a lot better. I made it async the following day and its even more fun. Most of peoples' problem is watching the AI type, it's not instant, but then not all (or even most) applications need to be instant. You can also avoid that by having it return everything at once instead of streaming style. Local absolutely can scale. All kinds of fun things can be done on a machine with 16GB of RAM, or 8GB if you work harder.
- DJHenk 3y ago> Most of peoples' problem is watching the AI type, it's not instant, but then not all (or even most) applications need to be instant. You can also avoid that by having it return everything at once instead of streaming style. Funny, for me it is the complete opposite. I created an interface in Matrix that does just that: return everything at once. But the lag annoys me more than the slow typing in the regular chat interface. The slow typing helps me keep me focused on the conversation. Without it, my mind starts wandering while it waits.
- klysm 3y agoNot as good as chatgpt 4 unfortunately, and they do have a moat. You could argue the most will fall in time but I’m not seeing chatgpt4 equivalents at the moment
- bredren 3y agoRunning an LLM locally and paying for access to OpenAI are two separate concerns. But to address both: is it very relevant what LLM you use right now? Local or hosted, openAI or other? It seems like the interface has converged around chat-based prompts. New ideas for tuning or improving the efficiency of foundational models are published almost every week. If one wants to build a product on top of of generative AI, why not simply start with what’s free or works with one’s dev environment? Presumably, the interaction with or API to text-based gen AI will be very similar no matter what engine is best for your use case at any given time. This would imply these backends will be swappable, the way web services are that copy AWS S3 APIs. So, to return to the point, can’t people just build their product with openAI or other and plan to move away based on the cost and fit for their circumstances? Couldn’t someone say prototype the entire product on some lower-quality LLM and occasionally pass requests to GPT4 to validate behavior? It seems far-fetched to believe this tech can be constrained by legislation. OpenAI can lobby all they want, it won’t necessarily buy them anything. Look what happened with FTX. Since LLMs can be run locally and the engines be black boxes to the user, how could a legislative act really prevent them from being everywhere—-especially given the public utility.
- ignoramous 3y ago> Couldn’t someone say prototype the entire product on some lower-quality LLM and occasionally pass requests to GPT4 to validate behavior? This, infact, might be a better way to do inference anyway: https://twitter.com/Francis_YAO_/status/1675967988925710338 https://twitter.com/Francis_YAO_/status/1675967988925710338 > So, to return to the point, can’t people just build their product with openAI or other and plan to move away based on the cost and fit for their circumstances? Depends. There are signs that folks are buying into GPT-specific APIs (like function calls) which may not be as easy to migrate away from.
- bredren 3y agoAsking because I have not implemented these yet: is there anything unique about the syntax that it can't just be copied?
- marinhero 3y agoMake a tutorial?
- two_in_one 3y agoI see no moral problems paying OpenAI for GPT Plus. it helps a lot in development. Their free speech-to-text 'whisper' is really good too. I'm going to use it + small local GPT for voice control. > I can currently run some scary smart and fast LLMs on a 5 year old laptop with no GPU. And, something useful or just playing? I played with local models, and will keep playing, training, experimenting. It's interesting, but not a solution, not yet.
- two_in_one 3y agoI'll take downvote as a sign you have nothing to say :) Just one warning, bad karma will be hard to fix.
- concordDance 3y ago> OpenAI has no moat, unless you give them money to write legislation. Their moat is that they had access to data sources which have since been clamped down on, eg reddit and twitter apis.
- travisjungroth 3y agoYou can still download Reddit archives with the same data they used.
- Obscurity4340 3y agoWhere can we aquire or access these local LLMs? How much space and specs does it actually require?
- dagaci 3y agohttps://gpt4all.io/index.html https://gpt4all.io/index.html is a good place to start, you can literally download one of the many recommended models. https://github.com/imartinez/privateGPT https://github.com/imartinez/privateGPT is great if you want do it with code.
- EagnaIonat 3y agoHuggingface has them all. https://huggingface.co https://huggingface.co
- lynx23 3y agoI tried, and decided it is not worth it. llama.cpp with a 13B model fit into RAM of my laptop, but pushes CPU temperature to 95 degrees within a few seconds, and mightily sucks the battery dry. Besides, the results were slow and rather useless. GPT is the first cloud application I deliberately use to push off computing and energy consumption to an external host which is clearly more capable of handling the request then my local hardware. I sympathize with the idea of wanting to run a local LLM, but IMO, this would require building a desktop with a GPU and plenty of horsepower + silent cooling and put it somewhere in a closet in my apartment. Running LLMs on my laptop is (to me) clearly a waste of my time and its battery/cooling.
- regularfry 3y agoSo I do actually want a really good games machine, and an AI worker box. Since I can't both use inference output and play games at the same time, having a ludicrously over-specced desktop for both uses actually makes sense to me.
- Jzush 3y agoI’m currently using the free tier ChatGPT web interface to help me with mundane coding tasks like JavaScript, php or css. Is there a local solution that is at least as intelligent as GPT 3.5 in that regard that I can run in a container?
- VladimirGolovin 3y agoCan you recommend some local LLMs that are (roughly) equivalent to ChatGPT?