8 ms·
Orca 2: Teaching Small Language Models How to Reason
- fgfm 3y agoOrca 2-13B consistently beat Llama 2-70B on most benchmarks in 0-shot. Hopefully, research papers will start to include Mistral/Zephyr 7B & Openchat 3.5. Even though they're smaller, they're getting competitive against much larger models and they're much cheaper to orchestrate.
- ple13 3y agoIt fails other benchmarks vs Mistral-7b. https://twitter.com/Teknium1/status/1726846755344634020 https://twitter.com/Teknium1/status/1726846755344634020 (There is some doubts about the validity of the comparaison in the comments)
- eurekin 3y agoAlso, worth mentioning the next tweet: Update, I benchmarked 13b Orca 2, its still not surpassing gpt4all score of Base Mistral or OpenHermes 2.5 7B: Hermes 2.5 7B Mistral score: 73.12% Mistral Base 7B score: 71.16% Orca 13B GPT4All score: 70.58% https://twitter.com/Teknium1/status/1726833004117635414 https://twitter.com/Teknium1/status/1726833004117635414
- davidkunz 3y agoFor smaller models, I'm impressed by Mistral-7b or fine-tuned variants like Zephyr. I use it regularly in Neovim[1] for mundane tasks (grammar correction, summaries, ...). I'm curious how Orca 2 performs, downloading it right now. [1]: with https://github.com/David-Kunz/gen.nvim https://github.com/David-Kunz/gen.nvim
- eurekin 3y agoI'd love to see some demo of that!
- GaggiX 3y agoAlso OpenChat-3.5v model (It has 7B parameters, I think it is also a Mistral finetuning), demo: https://openchat.team/ https://openchat.team/
- schleck8 3y agoNice, it passes the weather test. I always ask open source models what the weather is like and see wether it hallucinates my location and a forecast. A few months ago without exception all models I tried (even larger ones) would just make up a temperature. Now it replies as it should Cool! > what's the weather like today? > I'm sorry, but I can't provide real-time weather information. However, I can help you with general information about weather conditions and forecasting.
- nodja 3y agooh wow this model is kinda amazing, it passes my "creative" tests that only chatgpt 3.5 did decently well with, I've recently been disillusioned that open source has been moving the wrong way due to the focus on benchmarks, but this model seems to hit the spot in usefulness in more whacky prompts ("write X in the style of Y" kinda prompts)
- sorokod 3y agoAlways surprised how poorly these models do on the benchmarks they claim to do well. OpenChat has a benchmark radar diagram[1] but but often fails on actual samples. [1] https://github.com/imoneoi/openchat https://github.com/imoneoi/openchat
- titaniumtown 3y agoHaven't seen this neovim plugin before! I'm setting this up right now.
- intended 3y agoI really really want this to work. However at this point - benchmark success is about as effective as results from someone who has been “taught the test” If say… Merck wanted to use this same model to reason out a logistics issue, or apply it to some business problem at scale - you’d have to deal with hallucinations all over the place. The best analogy I have right now is that improved results on benchmarks are like better acting from Hugh Laurie as House. If you want to watch a show - great (generative work) If you want to get a prescription - then not so much.
- candiddevmike 3y agoI'm not a real AI doctor, I just play one on chat.openai.com.
- FFP999 3y agoAt the moment I read "how to reason" in the headline my bullshit detector started to go off. LLMs do not reason, they do not think, they are not AGI. They generate by regurgitating.
- coderaptor 3y agoI haven’t heard a definition of “reasoning” or “thinking” that proves humans aren’t doing exactly that same probabilistic regurgitation. I don’t think it’s possible to prove; feels like a philosophical question.
- btbuildem 3y agoAre we beginning to see "specialized SLMs"? We've already seen some pretend-agent based solutions (where the same model is given several different roles and made to act as eg. ceo / architect / dev / sales in a startup). I wonder if the way forward is to train smaller models with different sets of "skills" or "neural affinities". One for reasoning, one for summarization, one for math, one for code, etc - then combining them into full-fledged solutions. Perhaps smaller models can be "better" at their specific domains/tasks than the giant generalist models can be at any of them.
- worldsayshi 3y agoIsn't this the whole idea with Mixture Of Experts approach that is GPT-4 is using?
- htrp 3y agoIsn't MoE with switch transformers massively inefficiemt compared to being able to customize which LLMs you are using? I've seen a lot of agent swarm concepts in the smaller llm space that seem to provide some feedback that this is a viable avenue of research.
- esafak 3y agoIs GPT-4's MOE based on combining specialized models?
- hobofan 3y agoYes, I think that is the general trend. Have one model tuned for reasoning that decides a plan, based on which you invoke other models as tools (see e.g. the ReWOO paper[0]). If I had to guess, an approach like this is what powers the recent Custom GPT/Assistant API products (based on the lag between tool invocations I would guess that they also re-prompt for plan adjustments between every set of tool calls). Do that with a small model and hot-swap LORAs, and it should be possible to build a quite powerful local assistant on consumer hardware. [0]: https://arxiv.org/abs/2305.18323 https://arxiv.org/abs/2305.18323
- trash_cat 3y ago
- Philpax 3y agohttps://huggingface.co/microsoft/Orca-2-13b https://huggingface.co/microsoft/Orca-2-13b https://huggingface.co/microsoft/Orca-2-7b https://huggingface.co/microsoft/Orca-2-7b
- kromem 3y agoA really important nuance here is that they are building on top of Llama-2, the pretrained model, and not Llama-2-chat. I really think the entire field is doing a degree of damage with the chat fine tuning beyond what might be expected, because regularly part of that chat instruction is an emphasis on identification as a LLM. The problem with this is that nearly all of the training data it's performing next token prediction on is text generated by humans. So there's an inherent narrowing of the model scope with most of the fine tuning I've seen such that while pretrained models are harder to use, I regularly prefer them over chat models when both are available as even at similar temperatures the quality and variety of language is much improved in the pretrained over chat model. This fine tuning was only introducing bias towards logical step by step analysis and problem solving techniques, and the results are great. But I'm willing to bet that an identical fine tuning on top of the chat model would have been much worse on the evaluations - not just the compounding of a typical fine tuning loss of a few percent, but more like a double digit relative difference. It's quite frustrating that the anxiety over model safety is likely throwing out tens of millions of dollars worth of data in the pretrained model when only chat models are available for the SotA, and I hope in the future a lighter touch is taken on fine tuning the pretrained model and instead of focusing on safety inherent to the model it is just set behind a safety oriented discriminator or 'editor' which filters or modifies responses accordingly. I'd happily take a 2-3x increased API cost for a much more broadly capable and performant model with similar safety characteristics but without the handicaps that come with it. So while a lot of the gains here might be due to the fine tuning, I expect at least part is shrugging off the baggage of the chat/safety fine tuning as well. Even in the first detailed example, we can see that while Llama-2 goes off rambling later on, its statement of the relative knowledge of John vs Llama-2-chat is much more clear and connected between initial conditions and result particularly regarding theory of mind (i.e. "he assumed" vs the latter's "it must be in").
- kromem 3y agoAdding to this - it's really interesting the safety stuff that *is* in this paper. Such as: > We probe some of the categories where we see a larger difference (e.g., violent) and observe that Orca 2 tends to counter the harmful positions more often (which is penalized by the metric), while models that have gone through RLHF safety training tend to decline to respond more often (which is rewarded by the metric). Or the fact Orca 2 is less likely to extend hate speech than Llama-2-chat which theoretically went through safety fine tuning even though Orca 2 did not have any explicit safety fine tuning. Research over the past year has really demonstrated (a) just how impactful fine tuning can be - to the point of transmitting capabilities from larger models to smaller, and (b) that we're still clumsily wading through that process with only partial clarity on best practices as the foundational pretrained models get better and better at astounding rates.
- alecco 3y ago> Progressive Learning: We start with LLaMA-2-7B or LLaMA-2-13B checkpoint and finetune it on the train split of FLAN-v2 dataset for one epoch. Note that FLAN-v2 dataset contains both zero-shot and few-shot problems. We then train on 5 million ChatGPT data from Orca 1 for 3 epochs. Then we train on the combination of 1 million GPT-4 data from Orca 1 and Orca 2’s 817K data for 4 epochs. I think people are missing why they are comparing against Llama-2 13B/70B. They improved Llama-2 7B/13B and reach the level of a 5-10x larger model of the same base. This is huge. Models on HF. https://huggingface.co/papers/2311.11045 https://huggingface.co/papers/2311.11045
- schleck8 3y agoYeah, the 13b model outperforms the 70b Llama 2. Goes to show how much potential there is on the software optimization front as opposed to just scaling in size
- T-A 3y ago...and quantized ones from the usual suspect: https://huggingface.co/TheBloke/Orca-2-7B-GGUF https://huggingface.co/TheBloke/Orca-2-7B-GGUF https://huggingface.co/TheBloke/Orca-2-13B-GGUF https://huggingface.co/TheBloke/Orca-2-13B-GGUF The 7B Q5_K_M one is small enough to run on an 8GB consumer GPU.
- ganeshkrishnan 3y agoAll the 13B files seems to be quantized.
- jpdus 3y agoIt isn't. Compared to the original Orca model and method which spawned many of the current SotA OSS models, Orca 2 models seem to perform underwhelming, below outdated 13b models and below Mistral 7b base models (e.g. [1]; didn't test myself yet, ymmv). [1] https://twitter.com/abacaj/status/1727004543668625618?t=R_vVes4snXOjEEJ72uRnsQ&s=19 https://twitter.com/abacaj/status/1727004543668625618?t=R_vV...
- yujian 3y agoI'm not sure if I'm missing something from the paper, but are multi-billion parameter models getting called "small" language models now? And when did this paradigm shift happen?
- Chabsff 3y agoNowadays, small essentially means realistically useable on prosumer hardware.
- nathanfig 3y agoRelative term. In the world of LLMs, 7b is small.
- hmottestad 3y agoAll the llama models, including the 70B one can run on consumer hardware. You might be able to fit GPT-3 (175B) at Q4 or Q3 on a Mac Studio, but that's probably the limit for consumer hardware. At 4-bit a 7B model requires some 4GB of ram, so that should probably be possible to run on a phone, just not very fast.
- iandanforth 3y agoReleased under the MS Research License, so not OSI and non-commercial, for the curious. https://huggingface.co/microsoft/Orca-2-13b/blob/main/LICENSE https://huggingface.co/microsoft/Orca-2-13b/blob/main/LICENS...
- amelius 3y agoThis is why imho Microsoft is way cooler than Apple. They have tons of published research. In Apple, even speaking about your research with a friend may result in severe punishment.
- jjtheblunt 3y agoApple publishes too, search for it for example, but much less.
- amelius 3y agoMuch, much, less. They are definitely not in the same league.
- jug 3y agoThis sounds quite exciting! Like Mistral all over again, only more transparent, open, and major backing probably as Microsoft are looking to significantly reduce costs now that they're expanding AI wide across their platforms? The approach truly feels like a next step in LLM design.
- Yuvrajs 3y agoOfficial Orca-2 demo is available on huggingface Spaces now - https://huggingface.co/spaces/ari9dam/Orca-2-13B https://huggingface.co/spaces/ari9dam/Orca-2-13B