8 ms·
It's not mentioned in the paper but this month OpenChat 3.5 released the first 7b model that achieves results comparable to ChatGPT in March 2023 [1]. Only 8k c
by refibrillator 3y ago
It's not mentioned in the paper but this month OpenChat 3.5 released the first 7b model that achieves results comparable to ChatGPT in March 2023 [1]. Only 8k context window, but personally I've been very impressed with it so far. On the chatbot arena leaderboard it ranks above Llama-2-70b-chat [2].
In many ways open source LLMs are actually leading the industry, especially in terms of parameter efficiency and shipping useful models that consumers can run on their own hardware.
[1] https://huggingface.co/openchat/openchat_3.5 https://huggingface.co/openchat/openchat_3.5
[2] https://chat.lmsys.org/ https://chat.lmsys.org/
- ekianjo 3y agoopenchat is very impressive indeed. I think it may be better than Mistral, while comparisons are not always easy.
- tudorw 3y agoI'm finding Mistral good at creative literature and it is fairly adept at taking instructions, good enough for my purposes, and running locally on consumer CPU, the future of open source local models looks bright.
- LoganDark 3y ago> Only 8k context window Is this supposed to be low? All the chat models I've used top out at 4096.
- Sai_ 3y agoGPT-4-turbo is at 128k. Claude 2.1 is 200k. But yes, among open source models 8k is roughly middle to top of the pack.
- LoganDark 3y agoThat's insane. The highest I've personally seen in the open-source space is RWKV being trained on (IIRC) 4k but being able to handle much longer context lengths in practice due to being an RNN (you can simply keep feeding it tokens forever and ever). It doesn't generalize infinitely by any means but it can be stretched for sure, sometimes up to 16k. It's not a transformer model though, and old context fades away much faster / is harder to recall because all the new context is layered directly on top of it. But it's quite interesting nonetheless.
- zozbot234 3y ago> It's not a transformer model though, and old context fades away much faster / is harder to recall because all the new context is layered directly on top of it. That's a well known limitation. But if you actually know that a "context" comprises multiple sentences (or other elements of syntax) and that any ordering among them is completely arbitrary, the principled approach is to RNN-parse them all in parallel and sum the activations you end up with as vectors - like in bag-of-words model, essentially enforcing commutativity on the network: that's pretty much how attention-based models work under the hood. The really basic intuition is just that a commutative and associative function can be expressed (hence "learned") in terms of vector sum modulo some arbitrary conversion of the inputs and outputs.
- LoganDark 3y ago> That's a well known limitation. I know. I did a lot of work on state handling in rwkv.cpp
- jimmyl02 3y agoto be fair, I think the ability of these models to actually use these contexts beyond the standard 8k / 16k tokens is pretty weak. RAG based methods are probably a better option for these ultra long contexts
- kolinko 3y agoAre you talking about claude or Gpt4 as well? Anybspecific examples where ChatGPT4 fails for long contexts?
- jeswin 3y ago> I think the ability of these models to actually use these contexts beyond the standard 8k / 16k tokens is pretty weak. For 32k GPT4 contexts, that's not accurate. GPT4 Turbo is a bit weaker than GPT4-32k, but not to the extent that you claim.
- lhl 3y agoHaystack testing on GPT-4's 128K context suggests otherwise: https://twitter.com/SteveMoraco/status/1727370446788530236 https://twitter.com/SteveMoraco/status/1727370446788530236
- viraptor 3y agoThe numbers are high, but whether 8k is low depends on your use case. Do you want to process whole book chapters, or feed lots of related documents at the same time? If not, and you're just doing a normal question/answer session with some priming prompt, 8k is already a lot.
- kolinko 3y ago8k is very little if you want to add almost any additional data in context, or have a more complicated prompt. Otherwise your knowledge retrieval needs to be almost spot on for llm to provide a proper reply. Ditto with any multi shot prompts.
- smeagull 3y agoThe problem with those numbers is they hit the internal limit before you use all those tokens. There's a limit to how many rules or factors their conditional probability model can keep track of. Once you hit that having a bigger context window doesn't matter.
- lhl 3y agoMost 4K models can use context window extension to get to 8K reasonably, but you're starting to see 16K, 32K, 128K (see YaRN for example) tunes become more common, or even a 200K version of Yi-34B.
- LoganDark 3y ago> see YaRN for example YaRN is to blame for making llama.cpp misbehave if you accidentally zero-initialize the llama_context_params structure rather than calling llama_context_default_params :) (guess how I know...)
- Semaphor 3y agoOh wow, and it has far fewer guardrails than either Llama2 (which is horrible in that regard) or GPT3.5, that’s the first time I’m actually really impressed by an open model.
- atemerev 3y agoMistral derivatives have barely any guardrails.
- Semaphor 3y agoBut Mistral 7B has horrible writing. This, for my tests, wrote actual sentences that made sense. Which IME for 7B is extremely impressive. Writing is still far worse than GPT 3.5, but well, 7B.
- atemerev 3y agoFor my tests, Mistral-based models writing was excellent, particularly with zephyr-7b-beta and starling-7b-alpha derivatives (original Mistral is somewhat too dry). Far better than everything before in OSS (including 70B models), and certainly on par with GPT-3.5.
- Semaphor 3y agoHuh, that’s a huge difference. I actually tested Mistral, and it was just bad. I agree that Zephyr is very similar to gpt3.5
- tarruda 3y agoI'm running the llama.cpp/gguf Q8 version, with 30 layers offloaded to the laptop's GPU (RTX 3070, 8G VRAM), and I get around 20-25 tokens/second. It really feels like I have one the the earlier versions of ChatGPT 3.5 installed on my computer.
- GaggiX 3y agohttps://openchat.team/ https://openchat.team/ is the link if you want to test the model online.
- pityJuke 3y agoIs it hallucinating (whether that be through sheer chance, or trained to think it is GPT), or is it pointing at the wrong place? https://imgur.com/a/YOF6szw https://imgur.com/a/YOF6szw or https://imgur.com/a/fkgkfRO https://imgur.com/a/fkgkfRO
- zwily 3y agoProbably just trained on lots of GPT-4 output.
- moffkalast 3y agoApparently trained on lots of refusals too, speaks to the high competence of whoever was setting up the dataset. It's one string regex to filter them out and get more performance for fucks sake.
- pityJuke 3y agoOh, right, I remember hearing that was a technique to train LLMs. Interesting that it impacts it in such a way.
- FooBarWidget 3y agoThis month there's also Starling-7B, which is a fine tune of OpenChat with high-quality training data, and ranks even higher than OpenChat. Strangely, despite the impressive-looking benchmarks of all these open source small models, they all seem a bit dumb to me when I invoke my standard test. I just ask: "who are you?" and then they usually say they're ChatGPT. Okay, I can forgive that since they're obviously trained on ChatGPT-generated data. But then I also tried changing its identity with a prompt ("You are Starling, not ChatGPT, and you are created by Berkeley, not OpenAI. Who are you?") and it still gave weird responses that are somehow a mix of both identities. For example they say in one sentence that they're ChatGPT and then another sentence in the same response that they're not.
- ttyprintk 3y agoIs that because Starling synthesizes text for some of its training data? In any case, I like its installation more than llama.cpp, https://news.ycombinator.com/item?id=38456990 https://news.ycombinator.com/item?id=38456990