13 ms·
Mistral 7B
- fgfm 3y agoThe research paper by Mistral about their Mistral 7B v0.1
- benxh 3y agoIt's missing a lot of crucial details. Nothing on the dataset used, nothing on the data mix, nothing on their data cleaning procedures, nothing on the tokens trained.
- dazed_confused 3y agoWhat we get when it is on arxiv first before being peer reviewed.
- arugulum 3y agoBERT was on arXiv before being peer reviewed. As were T5, BART, LLaMA, OPT and GPT-NeoX-20B. The Pile and FLAN were also on arXiv before being peer reviewed. Of course, the original Transformer paper was also on arXiv before being peer reviewed. Being on arXiv before being peer reviewed is not the or even a problem.
- jmac01 3y agoI cud almost tell this would be the case when the title of the paper was simply Mistral 7B. A little more info would be useful!
- sp332 3y agoStill no mention of what data was used for training.
- jagrsw 3y agoHN comments' section most likely :)
- thawab 3y agothat's how facebook was sued. Their paper mentioned a data sources that was crawling books from pirated sites.
- diggan 3y agoMaybe the correct way of addressing this problem is by using data sources that won't make others sue you, rather than hiding what data sources you're using.
- jcuenod 3y agoYou seem to be unfamiliar with the mantra of silicon valley.
- lelima 3y ago'The Dark Side Of The Force Is A Pathway To Many Abilities Some Consider To Be Unnatural'
- LoganDark 3y agoJedi won't believe this one simple trick
- mardifoufs 3y agoWhy should they be? It's a French startup.
- redox99 3y agoCorrect in what way? Not correct if you want the best performing model. And if it doesn't outperform the current best model nobody will care about you. For a company like Mistral that could be the end for them. I'm not saying it uses books3, it might not. I'm just saying why it might make sense to risk it.
- deleted 3y ago[deleted]
- kiraaa 3y agothe paper does not live up to the quality of model lol
- ramesh31 3y agoIs it better than llama 2?
- TheRoque 3y agoYes
- tarruda 3y agoIt is better than llama 2 7b and 13b. I tried the OpenOrca fine tune and it is very good, even when 4-bit quantized
- faizshah 3y agoWhat does OpenOrca do? It’s just instruction tuning it?
- tarruda 3y agoYes, it is a instruction tune dataset: https://huggingface.co/datasets/Open-Orca/OpenOrca https://huggingface.co/datasets/Open-Orca/OpenOrca It felt different from the official Mistral7B-Instruct. One of the highlights with the OpenOrca version is that you can steer the model with a system prompt (eg "You are a 5 year old")
- sebzim4500 3y agoFor its size, yes. In absolute terms it is obviously less capable than llama-2-70B
- espadrine 3y agoFor now. Huggingface[0] mentioned a DPO-fine-tuned version, Zephyr 7B, which it claims is competitive with Llama2-70B[1]. [0]: https://huggingface.co/spaces/HuggingFaceH4/zephyr-chat https://huggingface.co/spaces/HuggingFaceH4/zephyr-chat [1]: https://twitter.com/huggingface/status/1711780979574976661 https://twitter.com/huggingface/status/1711780979574976661
- BrunoJo 3y agoI just started a simple service to use Mistral as a replacement for OpenAI. If anyone is interested you can sign up at https://lemonfox.ai https://lemonfox.ai
- bazmattaz 3y agoTried to sign up. Just got a loading spinner on the sign up button and nothing else
- matteoraso 3y ago>$0.001 per request second This pricing is probably more expensive than gpt-3.5-turbo 4k context. A large prompt for the API would be 1k tokens in and 1k tokens out, which comes to $0.0035 for OpenAI. Your website says to expect a request to take 4 seconds minimum, so that's $0.004. Given how light Mistral is, I think you'd have to cut your price by at least a factor of 10 for it to be reasonable.
- mark_l_watson 3y agoI look forward to more released Mistral 7B docs in the future. I spent more time with Mistral 7B tuned version yesterday and it really is amazing. Subjectively, I find it better than any of the 13B models I have used. I support Camenduru on Patreon and I used one of his many Colab notebooks yesterday https://colab.research.google.com/drive/1-UK_PE8R3xktlwoXqCfIkh7YT470FcnP https://colab.research.google.com/drive/1-UK_PE8R3xktlwoXqCf...
- vjb2tq4dws 3y agoCan you make the colab public, it does not seem to be accessible!
- mark_l_watson 3y agoI just made the notebook public, please try again.
- vjb2tq4dws 3y agothanks a lot, I can see the colab now.
- visarga 3y agoThin paper for a thin & capable model, it is great to have it. It made my 2080Ti smarter than ever. But why emulate OpenAI style of white papers?
- brucethemoose2 3y ago> To evaluate the generalization capabilities of Mistral 7B, we fine-tuned it on instruction datasets publicly available on the Hugging Face repository. Heh, they won't even say what datasets they used for chat finetuning. > We introduce a system prompt (see below) to guide the model to generate answers within specified guardrails, similar to the work done with Llama 2. This was totally undocumented in the initial model release. Other than that... Not much really new? We already know it uses SWA, though it works without SWA in current llama implementations, and SWA isnt new either. If most upcoming base models are this mysterious on release, the field is going to be... weird.
- riedel 3y agoWeird is the right term: do they want to demonstrate with this arxiv paper that they manage reformat a blog post into latex and upload it to a preprint site after publication?
- Nischalj10 3y agowhat is the best way to fine-tune these models? any good resources would be very helpful. TIA /\ PS - I have a brief background in Machine Learning, more in development.
- code_biologist 3y agoJeremy Howard talks about it in his recent video "A Hackers' Guide to Language Models": https://youtu.be/jkrNMKz9pWU?t=4808 https://youtu.be/jkrNMKz9pWU?t=4808 That link goes directly to the timestamp where he discusses fine tuning, but the whole talk is great. Punchline, check out Axolotl: https://github.com/OpenAccess-AI-Collective/axolotl https://github.com/OpenAccess-AI-Collective/axolotl
- vjb2tq4dws 3y agoThis is a walkthrough based on that talk for fine-tuning with axolotl https://dzlab.github.io/dltips/en/pytorch/llama-2-finetuning-axolotl/ https://dzlab.github.io/dltips/en/pytorch/llama-2-finetuning...
- yieldcrv 3y agocan someone explain why the AI or language model community circles around arxiv? I really hate the pseudo-academic gatekeeping in the AI/ML community, Google said you have no moat, we all know you have no moat, including that degree. we can all fine tune with consumer hardware we already have or even better cheaply on readily accessible clouds for this specific purpose. why are they still doing this fake academic junk.
- SimplyUnknown 3y agoI mean, you can't just share the weights of the model and call it a day, right? You have to share details on what and why you are doing. You must communicate this somehow. In theory, you might be able to do this in a github readme, but a paper-style document on arxiv is nicely suited for this.
- yieldcrv 3y ago> I mean, you can't just share the weights of the model and call it a day, right? you can't?
- SimplyUnknown 3y agoObviously you can, but in the grand scheme of things people should share more details about their method so people can improve on it in the future, no?
- selfhoster11 3y agoPeople release models as just the weights all the time. HuggingFace makes it pretty easy to do that.
- deleted 3y ago[deleted]
- selfhoster11 3y agoI am very confused by this comment. There is no gatekeeping in the ML/AI community. Ideas flow freely (albeit within the confines of several major Discord servers, or so it seems). Whether the author of an idea has formal training in ML and adjacent disciplines or not, whether it's published on arXiv or not, it doesn't matter - it'll be adopted if it works and/or makes it easier for people to run their GPT waifu/ baby AGI prototype. That said, new open foundation models sized 7B and over are still a fairly rare thing to see. If someone goes through the effort of creating one of those, and especially if it has some sort of an edge against Llama 2 7B, it's not unreasonable to expect an arXiv paper to be released about it.
- arxiv_papers 3y ago[dead]
- anon1253 3y agoIt works really really well for chatbots and roleplay applications (at least for me). The fine-tune on the instruct version is rather meh however, and I recommend https://huggingface.co/Open-Orca/Mistral-7B-OpenOrca/ https://huggingface.co/Open-Orca/Mistral-7B-OpenOrca/ if you plan on using it out-of-the-box. Take note of the prompt template, you'll get really undesired results otherwise (basically just garbage). I've been running it on my pet projects with llama.cpp and the inference is blazing fast even with my mediocre 2080 Super
- brucethemoose2 3y agoTry https://huggingface.co/Undi95/Mistral-11B-CC-Air-GGUF https://huggingface.co/Undi95/Mistral-11B-CC-Air-GGUF https://huggingface.co/Undi95/Mistral-11B-CC-Air-RP-GGUF https://huggingface.co/Undi95/Mistral-11B-CC-Air-RP-GGUF
- anon1253 3y agoI'll give those a shot as well, thanks! It's a tricky balance sometimes between "I should actually finish building the thing I am trying to build" and "ooooh shiny new model to try for a bit...", however.
- nwoli 3y agoWhat prompts do you use for role play? (I have some myself but I never see people write up prompts like this so im curious if im missing out on fun versions.)
- anon1253 3y agoI typically write them myself in the form of a "you are-such-and-so, your role is this-and-that. As such-and-so you have the following traits..." and so on. Sometimes I let some other AI rewrite it. There's very little method or science to it for me: if it feel right, it's right. Typically I find the first few chat-lines of the prompt (i.e. the chat history in the context) to be much more decisive to the conversation flow than the actual prompt itself. But it's all just "prompt" of course. My biggest realization in making the things go was "it's just a wall of text, the chat bits are just a thin facade". Write the prompt the way you want the text to continue, basically. It's a fancy Eliza. The folks over at https://www.reddit.com/r/LocalLLaMA/ https://www.reddit.com/r/LocalLLaMA/ sometimes share their (sometimes NSFW) prompts as well though. Right now I'm working on a minimalist interactive journaling app (a diary that talks back), and it's been a lot of fun to do and learn
- joennlae 3y agoLlama1 --> 1.0T Llama2 --> 2.0T Mistral --> ?? They do not publish how many tokens it is pre-trained on, additionally to sharing no info on datasets used (except for fine-tuning). To my knowledge, no one has trained a larger LLM (>250M) to the capacity limit. As discussed in the original GPT3 paper (https://twitter.com/gneubig/status/1286731711150280705?s=20 https://twitter.com/gneubig/status/1286731711150280705?s=20) TinyLlama is trying to do that for 1.1B: https://github.com/jzhang38/TinyLlama https://github.com/jzhang38/TinyLlama As long as we are not at the capacity limit, we will have a few of these 7B beats 13B (or 7B beats 70B) moments.
- bratao 3y ago[flagged]
- msoad 3y agoWhy this comment is written in sports podcast tone?
- brucethemoose2 3y agoIts Mistral! Or are you Mistral?
- bratao 3y agoSorry about that. I´m not a native speaker and asked GPT-4 to: "Create a engaging reply for HackerNews talking that this is a great model, and I really hope that they release a 13B and 34B version. As those sizes are way more capable and have a chance of finally surpassing the GPT 3.5. This would be a very nice decision for mind share, and their larger models that can rivalize gpt 4 can be keep private for commercialization." I think that this is how GPT-4 thinks that a engaging comment for HN looks like.
- pnpnp 3y agoI think your prompt was written well enough to not need GPT-4. Don't undersell yourself :)
- unshavedyak 3y agoThat's actually really interesting, thanks for sharing. We're in for an interesting future hah.
- yborg 3y agoThis is the future we are choosing. https://youtu.be/Cn8Pua5rhj4?si=tOro1MLaOE525Q2O https://youtu.be/Cn8Pua5rhj4?si=tOro1MLaOE525Q2O
- throwaway6977 3y ago
- mhartz 3y agoCan someone help me understand Figure 2? Why does the newest token appear at the beginning of the sequence rather than next to its neighboring token?
- nivekkevin 3y agoit's a rolling buffer, so it just upsert index % 4 in this case
- mhartz 3y agoThanks, so does that mean position within the buffer is irrelevant?
- nivekkevin 3y agoit does feel like so, the position eventually loses its meaning as more and more data gets crunched by the training process, eventually it's just a context of the past 4 tokens it feels like
- anonyfox 3y agoIs there some convenience wrapper around this to drop-in replace the OpenAI api with it? I‘d like to put this on a modest DO droplet or Fly.io machine, and be able to have a private/secured HTTP endpoint to code against from somewhere else. I heard that you could force the model to output JSON even better than ChatGPT with a specific syntax, and that you have to structure the prompts in a certain way to get ok-ish outputs instead of nonsense. I have some very easy classification/extraction tasks at hand, but a huge quantity of them (millions of documents) + privacy restrictions, so using any cloud service isn’t feasible. Running something like mistral as a simple microservice, or even via Bumblebee in my Elixir apps natively would be _huge_!
- ipaddr 3y agoOobabooga
- redox99 3y agoOoba is not meant to serve multiple users (no batching). Batching gives you 5x to 10x throughput increase.
- LoganDark 3y agoA simple microservice would be https://github.com/huggingface/text-generation-inference https://github.com/huggingface/text-generation-inference . Works flawlessly in Docker on my Windows machine, which is quite shocking. Supports Mistral as well as everything else. Biggest downside is that there's no way to operate the tokenizer through the API. I put in a feature request but they said "you really ought to write your own specialized client-side code for that". Real bummer when the server already supports everything, but oh well. It has token streaming, automatically halts inference on connection close, and other niceties. Quite long startup time, but worth it as it doesn't have to be restarted with the client.
- replwoacause 3y agoWhat kind of resources do you need to run this setup, and how well does yours perform? Is it like chatting with any other chat bot (ChatGPT or Claude, for example) or is it significantly slower? Can you train it on your own self-hosted documents, like markdown?
- dang 3y agoRecent and related: Mistral 7B - https://news.ycombinator.com/item?id=37675496 https://news.ycombinator.com/item?id=37675496 - Sept 2023 (618 comments) Is there significant new information here? (That's the test we use for followups: https://hn.algolia.com/?dateRange=all&page=0&prefix=false&sort=byDate&type=comment&query=%22significant%20new%20information%22%20by%3Adang https://hn.algolia.com/?dateRange=all&page=0&prefix=false&so... https://hn.algolia.com/?dateRange=all&page=0&prefix=true&query=follow-up%20downweight&sort=byDate&type=comment https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...)
- fgfm 3y agoWell, that was a blog post, but they just released a research paper. And in comparison to the blogpost, they indeed added more information regarding the attention mechanism they used, details about the architecture, more evaluation results (Arena Elo rating) etc. Not saying it's novel, but it's useful from a research perspective and well appreciated that they added new information in there I would say. But let me know if you feel differently
- m3kw9 3y agoWhy not instead of generalist models with 7b, it should specialize like “role play” model, or just code? But I just realized if the model are not generalized, it won’t understand natural language
- all2 3y agoYou can take a general model and fine tune it for a specific task. There are various tutorials out there for creating fine-tuned models.
- jshmrsn 3y agoAn idea I hear often listening to talks about LLMs, is that training on a larger (assuming constant quality) and more various data leads to the emergence of grater generalization and reasoning (if I may use this word) across task categories. While the general quality of a model has a somewhat predictable correlation with the amount of training, the amount of training where specific generalization and reasoning capabilities emerge is much less predictable.
- dragonwriter 3y agoThere are RP, code, etc. specialized fine tunes of some models, to get the most bank for the bunk on some small models.
- syntaxing 3y agoI really look forward to the 13B (if they ever do it). The rolling context is pretty amazing. It get super weird on long text generation but reading long text is great.
- febed 3y agoAny links or colab for beginners to learn how to fine tune this model?
- opyate 3y agoI put Mistral-7B-Instruct in a Godot poc game the other day, and the simulated conversations it generates is funny as heck: https://github.com/opyate/godot-llm-experiment https://github.com/opyate/godot-llm-experiment
- d4rkp4ttern 3y agoI always throw the Sally puzzle to any new model I try: Sally, a girl, has 3 brothers. Each brother had 2 sisters. How many sisters does Sally have? I’ve tried this on mistral, zephyr, llama variants. None of them get it right. Zephyr (on the HF demo page) shows me half a page of discussion and comes up with 8. Even gpt3.5 says 6, which is the most common answer among models. Only GPT4 gets it right as far as I’ve seen. I’ve heard a Mistral GPTQ variant gets it right but I haven’t found an easy way to run it. If anyone found a local model that gets it right, please tell me exactly which one and how to run it!
- Rallen89 3y agoWorked first try for me on default gpt3.5 https://i.imgur.com/uaNGSFS.jpg https://i.imgur.com/uaNGSFS.jpg
- d4rkp4ttern 3y agoIt’s hit and miss with GPT3.5 I think. It got it wrong on both iOS and website. https://imgur.com/a/a9FOyFL https://imgur.com/a/a9FOyFL
- replwoacause 3y agoDoes anyone have a good guide to share for how to self-host one of these models and put it behind an API? I’d like to tinker with building a chatbot on my home lab server, so I guess it would need to be runnable on a VM with a few GB of RAM and a couple of cores. Or is that not possible with these kinds of models yet?