4 ms·
If you're just looking to test it out, it's probably easiest to wait for llama.cpp to add support (https://github.com/ggerganov/llama.cpp/issues/6120 https://gi
by zone411 3y ago
If you're just looking to test it out, it's probably easiest to wait for llama.cpp to add support (https://github.com/ggerganov/llama.cpp/issues/6120 https://github.com/ggerganov/llama.cpp/issues/6120), and then you can run it slowly if you have enough RAM, or wait for one of the inference API providers like together.ai to add it. I'd like to add it to my NYT Connections benchmarks, and that's my plan (though it will require changing the prompt since it's a base model, not a chat/instruct model).
- logicchains 3y ago>it's probably easiest Cheapest maybe, but easiest is just to rent a p4de.24xlarge from AWS for a couple hours to test (at around $40/hour..).
- zone411 3y agoI'd expect more configuration issues in getting it to run on them than from a tested llama.cpp version, since this doesn't seem like a polished release. But maybe.
- v9v 3y agoThe NYT Connections benchmark sounds interesting, are the results available online?
- zone411 3y agoGPT-4 Turbo: 31.0 Claude 3 Opus: 27.3 Mistral Large: 17.7 Mistral Medium: 15.3 Gemini Pro 1.0: 14.2 Qwen 1.5 72B Chat: 10.7 Claude 3 Sonnet: 7.6 GPT-3.5 Turbo: 4.2 Mixtral 8x7B Instruct: 4.2 Llama 2 70B Chat: 3.5 Nous Hermes 2 Yi 34B: 1.5 The interesting part is the large improvement from medium to large models. Existing over-optimized benchmarks don't show this. - Max is 100. 267 puzzles, 3 prompts for each, uppercase and lowercase - Partial credit is given if the puzzle is not fully solved - There is only one attempt allowed per puzzle, 0-shot. - Humans get 4 attempts and a hint when they are one step away from solving a group I hoped to get the results of Gemini Advanced, Gemini Pro 1.5, and Grok and do a few-shot version before posting it on GitHub.
- stolsvik 3y agoWhere is this? I googled a bit, and found the game - but using it as a benchmark sounds genious!!