4 ms·
And this is only the beginning, I don't believe for a second that "AI" will belong to the big corporations. These LLM models have no benefit from running "in t
by tyfon 3y ago
And this is only the beginning, I don't believe for a second that "AI" will belong to the big corporations.
These LLM models have no benefit from running "in the cloud" except for processing power. Lots of disadvantages though, especially in data safety, "leaked chats to other users", privacy, bans etc.
- abraxas 3y agoWe really need to one of the two things happen. Either the models are somehow able to run in regular CPU DRAM or we see the GPU makers to finally give us sensible amounts of VRAM. It's a travesty that a card I bought in 2016 still has more VRAM than many of the flagships being sold today. This card is 7 years old for goodness sake!
- tyfon 3y agoI am really hoping we'll get something like AMD infinity fabric up and running with GPUs accessing system memory or something like that soon. Then there is also moore's law :)
- acapybara 3y agoMoore's law, sure. But rn it's totally doable to run a 65B or 100B model on CPU with a reasonable workstation. Does Moore keep us from wiring more memory onto a GPU? Or making a GPU with expandable memory (slots?)
- tyfon 3y agoI think the biggest issue with VRAM right now is that it's very expensive compared to regular system memory, perhaps with the exception of the DDR5 for the latest CPUs. But there should also be absolutely no issue in making a commercial TPU like google has internally for inference with more but less expensive ram and sell it. There surely must be a market now with these new models.
- Taek 3y agoI think part of the issue is that consumers have never needed more than 12 GB of VRAM before. Game developers just don't have requirements that high. Crypto mining also doesn't require that much. Now that there's clear demand in the hobbiest market for GPUs >100GB of vram, its more likely that manufacturers will step up with cheaper solutions.
- Der_Einzige 3y agoIt all comes down to the VRAM. The average person will never have the money to buy a single H100 96gb let alone a DGX server or it's future equivalents You get fundamentally more powers when you add more VRAM in ways that are just hard to explain to folks outside of this ecosystem. Everything around the VRAM are basically small details in comparison
- dalys 3y agoPlus, I assume people want to have their assistant on their phone, not their desktop computer. So until everything can run locally on your phone, I think people will prefer the cloud versions.
- dragonwriter 3y agoI assume some people want their assistant to work on data they don’t want to share with megacorps and governments, and some of those people can figure out how to make their phone talk securely to a home server over the internet.
- pmoriarty 3y ago"Everything around the VRAM are basically small details in comparison" ...with current algorithms and our lack of understanding and insight in to how/why they work on a deep level or what intelligence and consciousness is. With time hopefully all of these will improve and perhaps future AI's of good quality will be affordable to mere mortals.
- SanderNL 3y agoI know it’s not average person money, but the average person is not looking for hardware to run inference on LLMs. A100 seem in the €15k ballpark and H100 double that. Lot of money but I am actually surprised. A dedicated regular guy could buy this. I mean people buy cars and don’t really need them either. Again not saying it is a bargain, but it’s not billionaires only territory and that is good news (it’s early days!).
- emporas 3y agoWell, corporate organizations developing LLMs and open source LLMs are not mutually exclusive. LLMs more lightweight so as to run on more mainstream hardware, but not as capable as corporate ones, may still be very useful. Running the software on site has many advantages as you outlined above, by highlighting the disadvantages. The OpenAssistant was/is trained on well structured data from humans for exactly that purpose, for deep learning. In the past most LLMs were trained on unstructured internet data, and they performed well enough. But it was only when OpenAI used reinforcement learning that really the model started to shine. In my opinion well structured data as input to the machine, have a long way to go. More lightweight models, a lot more precise, a lot faster execution and a lot less memory usage are certainly possible. Most probably we are at the end of the road for the usefulness of structured data. I remember reading an article "Why Large Language models are over", meaning that smaller models but better trained, with better data and algorithms are the way to go.
- tyfon 3y agoNo they are not exclusive. But we do need alternatives to the big corp cloud models that are fully open source :)
- consumer451 3y ago> bans It feels extremely naive to think that all bans are a bad thing. Let's say that a criminal org starts a fully automated system to scam grandmas out of their savings. A cloud based service could ban them. A self-hosted system could not.
- lolc 3y agoThe same argument would work for a printer. Or nmap. It would be safer to not let people have these tools. Because criminals do make use of them too. Yet it is widely regarded as a good thing that nmap can be distributed and printers can be bought. Why are these models special?
- paulryanrogers 3y agoThey can fool the elderly by imitating their offspring, increasingly convincingly and at scale? Whereas nmap and printers cannot (yet)
- lolc 3y agoPrinters are often used to produce convincing fakes.
- welshwelsh 3y agoIf it was possible to guarantee that only felony-level criminal activity is banned, you might have a point. Realistically though and as we have seen with ChatGPT, if models can be censored they will be censored to the point where it affects normal people. Most people using chatbots have experienced "as an AI model, I can't do that" because of bullshit ethics. So... sorry about your savings Grandma, but I'm still going to fight for uncensored AI models. Fraud is already illegal, and if it happens we can prosecute the offenders.
- jstarfish 3y agoActually laughed at loud at "we can prosecute the offenders." We've prosecuted what, one (or two?) Nigerian princes in my lifetime. Grandma isn't leaving you anything when she passes if all of her savings were plundered by scammers while she was still alive. Preventing elder abuse is in your best interest.