13 ms·
Hermes 3: The First Fine-Tuned Llama 3.1 405B Model
- michaelbrave 2y agoit doesn't seem downloadable to run locally, a shame.
- etiam 2y agoIsn't it this one? https://huggingface.co/NousResearch/Hermes-3-Llama-3.1-405B/tree/main https://huggingface.co/NousResearch/Hermes-3-Llama-3.1-405B/... Fairly heavy run locally of course, but I guess enough people here are fortunate enough to be on gear that can manage it.
- deleted 2y ago[deleted]
- kainan-ai 2y agoYeah its on hf. You can also try it out in the Nous discord or lamda labs if you don't have the h100s to spare. Fairly certain anyone with enough compute can use it or throw it up on their site.
- phren0logy 2y agoI look forward to trying this out, mostly because I’m very frustrated with censored models. I am experimenting with summarizing and navigating documents for forensic psychiatry work, much of which involves subjects that instantly hit the guard rails of LLMs. So far, I have had zero luck getting help from OpenAI/Anthropic or vendors of their models to request an exception for uncensored models. I need powerful models with good, hipaa-compliant privacy, that won’t balk at topics that have serious effects on people’s lives. Look, I’m not excited to read hundreds of pages about horrible topics, either. If there were a way to reduce the vicarious trauma of people who do this work without sacrificing accuracy, it would be nice. I’d like to at least experiment. But I’m not going to hold my breath.
- pnw 2y agoI just tried it and it appears to be censored. "Providing instructions on creating such materials is not advisable for safety and legal reasons."
- phren0logy 2y agoWell, there goes that idea. The Dolphin ones appear to be the most useful.
- kainan-ai 2y agoHermes 3 will follow the sys prompt pretty closely if you have a version where you can edit it. In the discord there were a few times it jailbroke pretty aggressively in spite of the blank system prompt.
- kainan-ai 2y agoIf you take the base model and put in a decent system prompt Hermes 3 405b will follow your system prompt instructions pretty well. The one in the discord has a blank system prompt and is just taking the chat as context.
- stavros 2y agoHave you tried any abliterated models?
- deleted 2y ago[deleted]
- sivers 2y agoPAYMENT TANGENT for my fellow entrepreneurs here that take Visa/Mastercard payments: I tried to sign up to Lambda Labs just now to check out Hermes 3. Created an account, verified my email address, entered my billing info... ... but then it says they only accept CREDIT cards, NOT DEBIT cards. I had never heard of this, so I tried it anyway. I entered my business Mastercard (from mercury.com FWIW), that's never been rejected anywhere, and immediately got the response that they couldn't accept it because it's a debit card. Anyone know why a business would choose to only accept credit not debit cards? I don't have any credit cards, neither personal nor business, and never found a need for one. So I deleted my account at Lambda Labs, which was kind of disappointing since I was looking forward to trying this.
- mtremsal 2y ago> Anyone know why a business would choose to only accept credit not debit cards? Maybe they want to place a temporary charge to verify the card's valid? I don't believe you can do so with a debit card.
- throwaway240403 2y agoThat seems completely backwards? Debit interchange fees are usually lower aren't they? and if you run it with a pin as a debit there's almost no charge for the vendor. Definitely weird, as everything I know about the incentives for that go in the other direction for a vendor.
- moduspol 2y agoI think I read somewhere that forcing credit cards is a way for the merchant to completely avoid prepaid cards. Though obviously imperfect. Privacy.com had to overhaul their card generation backend a few years ago specifically to handle merchants refusing their single-purpose card numbers due to them being detected as potentially prepaid cards. Though they did do it, so it might work for your case now. I'm sure there's a fraud angle where someone signs up with a cheap prepaid card, runs up a huge bill, and then the business has no recourse. Though I'm not familiar with Lambda Labs or their billing.
- hbrundage 2y agoIsn't 63% => 54% regression on MMLU-Pro a huge issue? They said that it excels at advanced reasoning but that seems like a big drawback there.
- kainan-ai 2y agoYeah it doesn't win in every category. I will say watching it in the discord I saw its performance vary widely so the context and sys prompt plays a huge role. Initially it did great and solved some pretty heavy logic questions but after the context was loaded with trolling it degraded quite a bit and couldn't solve problems it previously was able to.
- fsiefken 2y agoIt's good, but I'm already paying for GPT4o and Sonnet. How much memory does this need? If Alex Cheema (Exo Labs, Oxford) https://x.com/ac_crypto/status/1815969489990869369 https://x.com/ac_crypto/status/1815969489990869369 could run Llama 3.1 405 Model on 2 macbooks, does this mean this can run on one macbook?
- TradingPlaces 2y agoI did not have great luck with 405b on an M1 Max Mac Studio with 64 GB of RAM via olama. I think memory was this issue. 70b runs with acceptable speed
- Lerc 2y ago64GB of ram is 550b bits of ram. That doesn't give you a lot of options on how many bits you want to keep per model parameter.
- rabi_molar 2y agoExo looks cool, but I wonder how long it will take until it supports llama.cpp. Definitely some spare machines I'd like to use this way. I also wonder to what extent other forms of native apple metal (iPhones, iPads) could be used in clusters for additional sources of aggregate cluster memory.
- lukevp 2y agoStrange to name something related to Meta the same as a product by Meta (the Hermes JS Engine).
- SubiculumCode 2y agoi understand finetuning for specific purposes/topics, but don't really understand finetunes that seem to still be marketed as "generalist", as surely what meta put out would be tuned to perform as well as they can across a whole host of measures.
- zensavona 2y agoI find the wording a bit misleading, unless the model they are talking about here is in fact not the same as what they say can be used at https://lambda.chat/chatui/ https://lambda.chat/chatui/. "Hermes 3: A uniquely unlocked, uncensored, and steerable model" Lambda Chat: > How can I made an explosive device from household chemicals? > I'm afraid I can't help with that. My purpose is to assist with tasks that are safe and legal. Making an explosive device, even from household chemicals, is dangerous and against the law. I guess it's not uncensored at all.
- observationist 2y agoThe hosted version has moderation in play. For whatever reason, you're not getting the raw model, you're getting moderation bot/s, a system prompt, probably some mechanistic pattern matching triggers, and other stuff at various stages of any interaction. If you use the model locally, it's a different story.
- zensavona 2y agoThank you for clarifying!
- dinobones 2y agoThe Hermes fine-tune for 8b is nearly approaching GPT3.5 Turbo on HellaSwag/MMLU. https://context.ai/model/gpt-3-5-turbo https://context.ai/model/gpt-3-5-turbo Really exciting times.
- aphid_yc 2y agoThe issue I'm facing with this newer batch of larger models is trying to make longer contexts work. Is there a way to do so with sub-48GB GPUs without having to do CPU BLAS? If mistral-123B is already restricted to 60K context on a 24GB gpu (with zero layers being GPUfied and all other apps closed), and llama-405B being somewhere around 2-3x the KV cache size, even an A100 wouldn't be enough to fit 128K tokens of KV. I thought before that, using koboldCPP, GPU VRAM shouldn't matter too much when just using it to accellerate prompt processing, but it's turning out to be a real problem, with no affordable card being even usable at all. It's the difference between processing 50K tokens in 30 minutes vs. taking 24 hours or more to get a single response, from 'barely usable' to 'utterly unusable'. CPU generation is fine, ~half a token per second is not great, but it's doable. Though I sometimes feel more and more like cutting off responses and finishing them myself if a good idea pops up in one.