4 ms·
Depends on the benchmark. Does well on other metrics when compared to open models. https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard https://hu
by binarymax 3y ago
Depends on the benchmark. Does well on other metrics when compared to open models.
https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...
- cosmojg 3y agoGiven that HellaSwag performance seems to correlate with reasoning ability more than other benchmarks, Falcon certainly look promising! Hopefully this is a clean result and not the product of dataset contamination.
- avereveard 3y agoI've given it a try, to having a chat is good, to follow langchain prompts it's not. I guess it depends on the type of work you want to extract from it.
- redox99 3y agoThat leaderboard is incorrect https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard/discussions/63 https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...
- lhl 3y agoActually it's HF's leaderboard that's bugged. Falcon is only on top since their MMLU scores are bugged across all LLaMA models: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard/discussions/63 https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...