6 ms·
Unfortunately I've found the current OSS models to be vastly inferior to the OpenAI models. Would love to see someone actually get close to what they can do wit
by kyledrake 3y ago
Unfortunately I've found the current OSS models to be vastly inferior to the OpenAI models. Would love to see someone actually get close to what they can do with GPT-3.5/4, except capable of running on commodity GPUs. What's the most impressive open model so far?
- asynchronous 3y agoProbably Wizard 30B, they released the training weights at every offset as well.
- bugglebeetle 3y agoIn my testing, the chat and instruct-tuned versions of MPT-30B is very close to 3.5 for many tasks, but of course the team who made it got bought up immediately and it’s licensed only for non-commercial use. I’m hoping the open source community runs with the base model in the same way they did with LLaMA.
- Roark66 3y agoHave you tried falcon 40b instruct? Also take into account that chatgpt likely has some preprompt and by talking to falcon or other OS models it's all in your hands. Furthermore, Not many people discuss the significance of proper output sampling. I myself used to just test open source models with the greedy decoding only. Who knows if they wouldn't even beat (not at all)OpenAI with some clever output sampling scheme.
- schappim 3y agoDoes anyone have a link / instructions on getting Falcon 40b to install on Apple Silicon? Apparently "Hugging Face" have some internal swift code that works (but it has not been released). I'm keen to see how it performs on a maxed out Mac Studio (with all that unified memory available).
- _ea1k 3y agoI thought this too, but it sounds like the performance of the GPU ultimately holds it back. Maybe it is the software, but I have yet to see a test on Mac silicon that performed well.
- schappim 3y agoIt is not too bad, 7B is doing 432 tokens per second on a Macbook[1]. [1] Video: https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/blog/147_falcon/falcon-7b.mp4 https://huggingface.co/datasets/huggingface/documentation-im...
- moffkalast 3y agoIdk, has anyone tried falcon yet? The support for running it remains nonexistent except for one fork of llama.cpp that isn't integrated into anything. This trend of every new model breaking compatibility really needs to stop.
- Roark66 3y agoI have and I am running it locally. Mostly in 7b variant which runs pretty much at chatgpt speed (when streaming) on my ryzen 3700 cpu + 32gb ram + 2xnvme in a mirror (it doesn't fit in the memory in its entirety, few gb go into the swap). Of course to run it like that I have to be running nothing else. No xorg, no chromium etc. Just a pure linux console. If you want to try falcon 40b instruct by yourself here I'd a public demo : https://huggingface.co/blog/falcon https://huggingface.co/blog/falcon Go to the bottom of the page.
- juliensalinas 3y agoLLaMA 30B or 60B can be very impressive when correctly prompted. Deploying the 60B version is a challenge though and you might need to apply 4-bit quantization with something like https://github.com/PanQiWei/AutoGPTQ https://github.com/PanQiWei/AutoGPTQ or https://github.com/qwopqwop200/GPTQ-for-LLaMa https://github.com/qwopqwop200/GPTQ-for-LLaMa . Then you can improve the inference speed by using https://github.com/turboderp/exllama https://github.com/turboderp/exllama . If you prefer to use an "instruct" model à la ChatGPT (i.e. that does not need few-shot learning to output good results) you can use something like this: https://huggingface.co/TheBloke/Wizard-Vicuna-30B-Uncensored-fp16 https://huggingface.co/TheBloke/Wizard-Vicuna-30B-Uncensored... The interesting thing with these Uncensored models is that they don't constantly answer that they cannot help you (which is what ChatGPT and GPT-4 are doing more and more).
- wiz21c 3y agoWhat are all the researchers in universities doing ? Couldn't they improve these models (they do have big brains after all) with tax payer's money and put the results under some cool open source license...
- dr_dshiv 3y agoThey are busily working on all the other large scale engineering projects that they are so good at producing and maintaining.
- albertzeyer 3y agoThey don't have enough money and thus computing resources. It's a huge gap.
- jona-f 3y agoYes, they are doing the improving, but then you need loads of money to do the learning no university can afford. So now big tech is hiring promising university researchers for good money to scale up their research. This could be solved by massive decentralization where millions of users provide compute with their gpus and i think it will be at some point, cause i believe foss is more powerful than this openai bs. There are people working on this, but afaik the techniques aren't quite there. You need a different kind of model with much more parallelization then what is currently used.
- wiz21c 3y agoI knew for the lack of money. But the parallelization idea, I didn't. Thanks for posting !
- DeathArrow 3y ago> There are people working on this, but afaik the techniques aren't quite there. You need a different kind of model with much more parallelization then what is currently used. What if crypto is switching from mindless hashing as proof-of-work to training AI models as proof-of-work? That would mean suddenly big computing resources are available.
- raydev 3y agoWhile OpenAI likely has some insights that open-source and closed-source competitors are lacking, OpenAI is mostly in the lead because they can burn absurd amounts of cash running an absurd amount of compute via their partnership with Microsoft.