4 ms·
Mistral-7B-OpenOrca. First 7B model to beat all other models <30B
- easygenes 3y agoThis new release is a finetune of Mistral-7B base model. They claim it achieves 98% of the eval performance of Llama2-70B.
- easygenes 3y agoHaving tested this model a lot today, I can say that it indeed in a lot of ways feels much like a 70B model, much to my surprise. One thing that really struck me: I asked it details about my local suburbs (minor cities in a larger metro). I have tried this on all different Llamas, and only the 70B models don't just hallucinate everything or tell me they don't know anything. In contrast, this model knows quite a bit about my local area. Not as good as Llama2-70B, but not incredibly far off.
- brucethemoose2 3y agoThough it does seem to he very coherent, I am still very suspicous of test data contamination in Mistral. XWin was the SOTA llama finetune last I checked, so we will see how it fares with Mistral...
- easygenes 3y agoXWin is good for AlpacaEval, but was optimized specifically for that. Evaluated more broadly, there are other better choices.
- squigz 3y agoCould you elaborate what you mean by "test data contamination" and why that's a specific concern for these models vs others?
- brucethemoose2 3y agoLLMs are evaluated by asking them a set of standard test questions and seeing how they respond. But they are not supposed to train on these test questions. If they do train on those test Q/A pairs, intentionally or not, the answers will be conspicuously good. A purposely seperate test dataset like that is standard in basically all ML research. But Mistral has given basically zero information about the Mistral 7B training dataset. And they have less reputation to lose from "cheating" than Meta does.
- squigz 3y agoInteresting. I've actually been thinking of that issue lately, since it seems an unavoidable problem (to my limited perspective on this field) - how do you avoid those test questions being spread around the internet and making their way into training data? Do you come up with new questions every time? Can you look for such data before training and remove it? Is there any reason to believe this issue is more inherent to Mistral, other than because Meta and the like have a larger reputation? Has Meta released the relevant information about their training data and evaluation?
- brucethemoose2 3y ago> How do you avoid those test questions being spread around the internet and making their way into training data? Purposely search for the big standardized ones in the training dataset and filter them out. If some scattered examples are reworded enough to be undetectable, thats probably not a big deal. > Is there any reason to believe this issue is more inherent to Mistral, other than because Meta and the like have a larger reputation? Its not any one specific smoking gun... - Even with Mistral's recent code dump, the release is still so light on detail that its kinda sketchy to me. - The base model beats llamav2 7B in metrics by a lot. - There is zero information on how to reproduce these metrics. This was also an issue for Llama, as testers had trouble reproducing their scores. - Mistral 7B doesn't seem to handle random questions completions as well as llamav2 7B base in my quick tests (though this is not true for the instruct model. Mistral 7B instruct "feels" quite intelligent). I am not a data scientist, I don't want to point fingers. But there are llama finetunes from AI startups with similar red flags, and one I briefly tested autocompletes hellaswag questions suspiciously well, even with some of the "incorrect" answers from that dataset: https://huggingface.co/datasets/Rowan/hellaswag/viewer/default/test https://huggingface.co/datasets/Rowan/hellaswag/viewer/defau... I have this sneaking suspicion that AI may have an "crypto moment" where everyone realizes lots of eval scores are fraudulent, and that significant parts of the training datasets are questionably legal.
- brianjking 3y agoAre there any quantized versions yet? https://huggingface.co/s3nh/Open-Orca-Mistral-7B-OpenOrca-GGUF/discussions https://huggingface.co/s3nh/Open-Orca-Mistral-7B-OpenOrca-GG... doesn't seem to have any weights uploaded. Secondly, is this instruction tuned?
- SushiHippie 3y agoGGUF: https://huggingface.co/TheBloke/Mistral-7B-OpenOrca-GGUF https://huggingface.co/TheBloke/Mistral-7B-OpenOrca-GGUF GPTQ: https://huggingface.co/TheBloke/Mistral-7B-OpenOrca-GPTQ https://huggingface.co/TheBloke/Mistral-7B-OpenOrca-GPTQ AWQ: https://huggingface.co/TheBloke/Mistral-7B-OpenOrca-AWQ https://huggingface.co/TheBloke/Mistral-7B-OpenOrca-AWQ Thebloke is really really fast, he quantized and uploaded them already yesterday.