3 ms·
OpenAI offering 128k context is very appealing, however. I tried some Mistral variants with larger context windows, and had very poor results… the model would
by coder543 3y ago
OpenAI offering 128k context is very appealing, however.
I tried some Mistral variants with larger context windows, and had very poor results… the model would often offer either an empty completion or a nonsensical completion, even though the content fit comfortably within the context window, and I was placing a direct question either at the beginning or end, and either with or without an explanation of the task and the content. Large contexts just felt broken. There are so many ways that we are more than “two weeks” from the open source solutions matching what OpenAI offers.
And that’s to say nothing of how far behind these smaller models are in terms of accuracy or instruction following.
For now, 6-12 months behind also isn’t good enough. In the uncertain case that this stays true, then a year from now the open models could be perfectly adequate for many use cases… but it’s very hard to predict the progression of these technologies.
- pclmulqdq 3y agoComparing a 7B parameter model to a 1.8T parameter model is kind of silly. Of course it's behind on accuracy, but it also takes 1% of the resources.
- coder543 3y agoThe person I replied to had decided to compare Mistral to what was launched, so I went along with their comparison and showed how I have been unsatisfied with it. But, these open models can certainly be fun to play with. Regardless, where did you find 1.8T for GPT-4 Turbo? The Turbo model is the one with the 128K context size, and the Turbo models tend to have a much lower parameter count from what people can tell. Nobody outside of OpenAI even knows how many parameters regular GPT-4 has. 1.8T is one of several guesses I have seen people make, but the guesses vary significantly. I’m also not convinced that parameter counts are everything, as your comment clearly implies, or that chinchilla scaling is fully understood. More research seems required to find the right balance: https://espadrine.github.io/blog/posts/chinchilla-s-death.html https://espadrine.github.io/blog/posts/chinchilla-s-death.ht...
- danielmarkbruce 3y agoIt's an order of magnitude comparison. Let's just agree it's 100x-300x more parameters, and let's assume the open ai folks are pretty smart and have a sense for the optimal number of tokens to train on.
- razodactyl 3y agoThis definitely. Andrej Karpathy himself mentions tuned weight initialisation in one of his lectures. The TinyGPT code he wrote goes through it. Additionally explanations for the raw mathematics of log likelihoods and their loss ballparks. Interesting low-level stuff. These researchers are the best of the best working for the company that can afford them working on the best models available.
- razodactyl 3y agoNah, it's training quality and context saturation. Grab an 8K context model, tweak some internals and try to pass 32K context into it - it's still an 8K model and will go glitchy beyond 8K unless it's trained at higher context lengths. Anthropic for example talk about the model's ability to spot words in the entire Great Gatsby novel loaded into context. It's a hint to how the model is trained. Parameter counts are a unified metric, what seems to be important is embedding dimensionality to transfer information through the layers - and the layers themselves to both store and process the nuance of information.
- visarga 3y agoUsually the 7B model is fine-tuned with "enriched" data, "textbook quality" generations from the 1.8T model. Riding on its coat tails.
- tannhaeuser 3y agoThat's my take-away from limited attempts to get Code Llama2 Instruct to implement a moderately complex spec as well, using special INST and SYS tokens even or just pasting some spec text along in a 12k context when Code Llama2 supposedly can honor up to 100k tokens. And I don't even know how to combine code infilling with an elaborate spec text exceeding the volume of what normally goes into code comments. Is ChatGPT 4 really any better?
- fancy_pantser 3y agoOpen researchers are trying to shrink and speed up 138K models e.g. YaRN https://github.com/jquesnelle/yarn https://github.com/jquesnelle/yarn It's very compelling and opens up a lot of use cases, so I've been keeping an eye out for advancements. However, inferencing on 4xA100s would be the target today for YaRN and 128K to get a reasonable token rate on their version of Mistral.