7 ms·
https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb... It doesn't include a compariso
by pocketarc 3y ago
https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...
It doesn't include a comparison to the paid ones, but it's a great place to look at different ones and see which ones you should experiment with and sink time into.
- tmikaeld 3y agoFound it a minute after i posted ;-) Have you tried any of the high-scored ones?
- garciasn 3y agoThey change somewhat frequently. We were previously working with some of the highest ranked that we could handle and we were getting acceptable results for the parameters of our tests. That said, if you’re looking for these to be a similar quality to OpenAI’s ChatGPT 4, they’re not even remotely close.
- tmikaeld 3y ago> if you’re looking for these to be a similar quality to OpenAI’s ChatGPT 4, they’re not even remotely close. I figured as much, but it's incredibly hard to find any comparison that's kept up to date. What I was after was a local javascript LLM coder that i could train on a local codebase, but even that was fruitless except for some unmaintained "promptr" project. I guess OpenAI embeddings is the only option.
- politelemon 3y agoI have tried Stable Beluga which was released recently. Even at 7B, the smallest, it does a decent job.
- pocketarc 3y agoI've been playing a lot with Llama 2 13B recently, and it's really not bad at all. With oobabooga[1] you get a proper UI for it and even get an OpenAI-compatible API, so you just change the endpoint in your OpenAI library and it all works. I've been using that to test changes to my bots. As another poster mentioned though, it's nowhere near the level of GPT-4. It's close enough to GPT 3.5 though, you should try it out! [1]: https://github.com/oobabooga/text-generation-webui https://github.com/oobabooga/text-generation-webui
- tmikaeld 3y agoThanks, I've had oogabooga bookmarked for a while, it's time!
- optimalsolver 3y agoWhat happens when the average approaches 100%? i.e., do these benchmarks taken together capture anything resembling general intelligence?
- CuriouslyC 3y agoWhen the benchmarks approach 100% it'll be time to create harder benchmarks. We'll know we're getting close to general intelligence when it's a major challenge for humans to create harder benchmarks.
- gpderetta 3y agoyou trigger the on-site atomics to prevent a containment breach.