3 ms·
Hi! We ran LSH filtering over datasets to remove all code that can be similar to HumanEval samples.
by mityamitya 3y ago
Hi! We ran LSH filtering over datasets to remove all code that can be similar to HumanEval samples.
- riku_iki 3y agoso, we have to trust your procedure..
- JegernOUTT 3y agoIt can be checked if the model predicts canonical solutions from humaneval. I understand it is not ideal, but at least you can check it yourself There are a bunch of other benchmarks too, check out the page https://huggingface.co/smallcloudai/Refact-1_6B-fim https://huggingface.co/smallcloudai/Refact-1_6B-fim Also, feel free to run any new benchmarks