3 ms·
Firstly, I'll say that it's always exciting to see more weight-available models. However, I don't particularly like that benchmark table. I saw the HumanEval s
by coder543 2y ago
Firstly, I'll say that it's always exciting to see more weight-available models.
However, I don't particularly like that benchmark table. I saw the HumanEval score for Llama 3 70B and immediately said "nope, that's not right". It claims Llama 3 70B scored only 45.7. Llama 3 70B Instruct[0] scored 81.7, not even in the same ballpark.
It turns out that the Qwen team didn't benchmark the chat/instruct versions of the model on virtually any of the benchmarks. Why did they only do those benchmarks for the base models?
It makes it very hard to draw any useful conclusions from this release, since most people would be using the chat-tuned model for the things those base model benchmarks are measuring.
My previous experience with Qwen releases is that the models also have a habit of randomly switching to Chinese for a few words. I wonder if this model is better at responding to English questions with an English response? Maybe we need a benchmark for how well an LLM sticks to responding in the same language as the question, across a range of different languages.
[0]: https://scontent-atl3-1.xx.fbcdn.net/v/t39.2365-6/438037375_405784438908376_6082258861354187544_n.png?_nc_cat=106&ccb=1-7&_nc_sid=e280be&_nc_ohc=7AIeH58EojUAb7BXK2c&_nc_ht=scontent-atl3-1.xx&oh=00_AfAs3tOUPoHfB2vkPKCZRfAhjDP0aOZH7SrYqgqVFGJWTg&oe=6645F5CA https://scontent-atl3-1.xx.fbcdn.net/v/t39.2365-6/438037375_...
- d3m0t3p 2y agoThat's funny you mentioned switching to another language, I recently asked chatGPT "translate this: <random german sentence>" And it translated the sentence in french, while I was speaking with it in english"
- coder543 2y ago[deleted]
- wongarsu 2y agoI see the science fiction meme of AI giving sassy, technically correct but useless answers is grounded in truth.
- coder543 2y agoBy ChatGPT, do you mean ChatGPT-3.5 or ChatGPT-4? No one should be using ChatGPT-3.5 in an interactive chat session at this point, and I wish OpenAI would recognize that their free ChatGPT-3.5 service seems like it is more harmful to ChatGPT-4 and OpenAI's reputation than it is helpful, just due to how unimpressive ChatGPT-3.5 is compared to the rest of the industry. You're much better off using Google's free Gemini or Meta's Llama-3-powered chat site or just about anything else at this point, if you're unwilling to pay for ChatGPT-4. I am skeptical that ChatGPT-4 would have done what you described, based on my own extensive experience with it.
- wesleyyue 2y agohumaneval is generally a very poor benchmark imo and I hate that it's become the default "code" benchmark in any model release. I find it more useful to just look at MMLU as a ballmark of model ability and then just vibe checking it myself on code. source: I'm hacking on a high performance coding copilot (https://double.bot/ https://double.bot/) and play with a lot of different models for coding. Also adding Qwen 110b now so I can vibe check it. :)
- andai 2y agoDidn't Microsoft use HumanEval as the basis for developing Phi? If so I'd say it works well enough! (At least Phi 3, haven't tested the others much.) Though their training set is proprietary, it can be leaked by talking with Phi 1_5 about pretty much anything. It just randomly starts outputting the proprietary training data.
- kristianp 2y agoHumaneval was developed for codex I believe: https://arxiv.org/abs/2107.03374 https://arxiv.org/abs/2107.03374
- coder543 2y agoI agree HumanEval isn't great, but I've found that it is better than not having anything. Maybe we'll get better benchmarks someday. What would make "Double" higher performance than any other hosted system?
- cosmojg 2y ago> My previous experience with Qwen releases is that the models also have a habit of randomly switching to Chinese for a few words. I wonder if this model is better at responding to English questions with an English response? Maybe we need a benchmark for how well an LLM sticks to responding in the same language as the question, across a range of different languages. This is trivially resolved with a properly configured sampler/grammar. These LLMs output a probability distribution of likely next tokens, not single tokens. If you're not willing to write your own code, you can get around this issue with llama.cpp, for example, using `--grammar "root ::= [^一-鿿ぁ-ゟァ-ヿ가-힣]*"` which will exclude CJK from sampled output.
- lhl 2y agoI'd recommend those looking for local coding models to go for code-specific tunes. See the EvalPlus leaderboard (HumanEval+ and MBPP+): https://evalplus.github.io/leaderboard.html https://evalplus.github.io/leaderboard.html For those looking for less contamination, the LiveCodeBench leaderboard is also good: https://livecodebench.github.io/leaderboard.html https://livecodebench.github.io/leaderboard.html I did my own testing on the 110B demo and didn't notice any cross-lingual issues (which I've seen with the smaller and past Qwen models), but for my personal testing, while the 110B is significantly better than the 72B, it doesn't punch above its weight (and doesn't perform close to Llama 3 70B Instruct from my testing). https://docs.google.com/spreadsheets/d/e/2PACX-1vRxvmb6227Au4HL0ROH5uxaFPJReCXbG1IU-uMj06d3VgYO5Mu2h_gvf3Fi3gdOT3mXEev3xDf-GkhW/pubhtml https://docs.google.com/spreadsheets/d/e/2PACX-1vRxvmb6227Au...
- justinlin610 2y agono this is different. it is for the base model. this is why i explain in my tweet that we just say for the base model quality we might be comparable. for instruct model, there is much room to improve especially on human eval. i admit that the code switching is a serious problem of ours cuz it really affects the user experience of english users. but we find that it is hard for a multilingual model to get rid of this feature. we'll try to fix it in qwen2.