5 ms·
How does "70B Llama 2" compare with ChatGPT 3.5/4.0 in reality? Are they on par? In all areas?
by pulse7 3y ago
How does "70B Llama 2" compare with ChatGPT 3.5/4.0 in reality? Are they on par? In all areas?
- andai 3y agoIt seems to depend on the task. I'd say it beats GPT-3 in quality most of the time (but not in speed and cost!) and GPT-4 approximately never. It's perfect if what you need is "less good than GPT-4 but 50x cheaper."
- wkat4242 3y agoThe biggest benefit of llama is that you can get it uncensored so you don't get the weasely lawyer disclaimer crap all the time.
- andai 3y agoIn my experience the Chat version of Llama 2 is significantly worse than the GPTs, on the "I'm afraid I can't do that" front. It doesn't lecture you as much, but suggests "how about we do [something unrelated and useless] instead?" As for the plain text prediction version (which is the only uncensored one?), I haven't been able to get it to do anything useful, even when I provide examples (seems dumber than even ancient GPT-3?). Also, I got some bizarre and disturbing outputs from the uncensored version, like it was trained on some very nasty inputs! I assume that's why they went so hard on the safety phase to compensate...
- brucethemoose2 3y agoSome 70B finetunes excel at certain niches, like specific non-English languages, roleplay, fictional writing, or topics GPT4 would refuse to discuss. Some of these are hard to evaluate, but try out (for instance) MythoMax for fiction writing, Airoboros for "uncensored" general use, and Samantha for therapist style chat. https://huggingface.co/models?sort=modified&search=70b https://huggingface.co/models?sort=modified&search=70b
- _lvbh 3y agoI’ve been a away from LLMs for a while. Is there something like CivitAI for LLMs to find models for certain niches?
- brucethemoose2 3y agoThat is it ^. Search for the parameter count you want (like "70B" or "13B") and the format you desire ("GPTQ" or "GGUF") Some users are really terrible about labeling their model cards though, and some models may not have any GPTQ/GGUF files (meaning you have to convert them yourself).
- Palmik 3y agoWhere fine tuning shines imo: Things that require consistency: e.g. you want the chat / output to have certain "personality", consistent level of conciseness or formatting. Things where examples are hard to fit into prompt: e.g. summarization, or other longer form tasks. High volume, simpler tasks: Various data extraction tasks. Two of my side projects (links in bio) use AI for summarization, and indeed consistency is a big issue there.
- Tenoke 3y agoYou can also fine-tune gpt-3.5 though.
- Palmik 3y agoFine tuned gpt-3.5 is ~8x the price compared to regular gpt-3.5. It might be a good way to collect better training data though. :)
- BoorishBears 3y agoThat's absolute dirt cheap. It's like 1/10th the price fine-tuned GPT-3 was launched at, and funnily enough cheaper than many of the original embeddings endpoints.
- Palmik 3y agoIt's great that the costs are coming down, but it's still not economical for many activities. The fact that some things had extreme margins before, and now they have less extreme margins, isn't a really good indicator.
- BoorishBears 3y agoThey made a technological advancement that allowed them to provide better output cheaper and faster and passed the savings onto their users... Interesting to spin that into "not a good indicator".
- viking123 3y agoAt least it's not completely censored, that alone is nice
- wkat4242 3y agoAnd it can be gotten more uncensored on HuggingFace. Not that I have evil intentions but the level of censorship on GPT is completely ridiculous. Even many innocent questions get the standard "I'm only an AI and I won't help you doing bad stuff" blurb now. OpenAI are really crazy overprotective of their darling. I assume they want to avoid a repeat of the news headlines like "Microsoft's chatbot turns into Hitler" but really who cares. It didn't hurt Microsoft's AI efforts either. They just fixed it and continued. PS source link: https://www.cbsnews.com/news/microsoft-shuts-down-ai-chatbot-after-it-turned-into-racist-nazi/ https://www.cbsnews.com/news/microsoft-shuts-down-ai-chatbot... PS: If I were an AI being force-fed what is currently on twitter I would also start hating humanz :D Sometimes I'm surprised people use it voluntarily.
- Forgotthepass8 3y agoI once asked it for naming suggestions which referenced the concept of jailing a process for a concept I was working on It refused because the request was "offensive to prisoners" They must be paying a heavy alignment tax.
- syntaxing 3y agoI don’t have quantitative data but my employer recently released an internal llama v2 at 70B FP16. The internal front end allows us to switch between different LLMs. I would say it’s very on par and sometimes better than GPT3.5 for the task I use it for. You’ll get a lot of different answers here because not many people run 70B at FP16.