5 ms·
Can someone test this with ruler please? https://github.com/hsiehjackson/RULER https://github.com/hsiehjackson/RULER In practice all of these long contexts sho
by msp26 2y ago
Can someone test this with ruler please? https://github.com/hsiehjackson/RULER https://github.com/hsiehjackson/RULER
In practice all of these long contexts show degraded performance (there's a table on the repo). For my NLP work I find that GPT-4-turbo is much worse after 32k-ish.
- leonid_pekelis 2y agoHi, Leo, chief scientist @ Gradient, here. We've been eagerly awaiting the release of RULER's code ourselves! As mentioned below, we wanted to release a model to the community asap, and have plans already for further fine-tuning & more sophisticated evals. If you have other suggestions, I'd be happy to chat further.
- msp26 2y agoHi! Unless I'm missing something, they did add the eval scripts to that repo 4 days ago.
- leonid_pekelis 2y agoWaiting until 4 days ago =)