4 ms·
This is quite interesting because I've specifically tried this kind of basic ensembling for my NYT Connections benchmark and it didn't work. This is something e
by zone411 3y ago
This is quite interesting because I've specifically tried this kind of basic ensembling for my NYT Connections benchmark and it didn't work. This is something everybody would try first before more complicated multi-step prompting, and yet since ChatGPT 3.5 I'm not aware of any papers showing that it works. It will be interesting to reproduce this result and learn more about how they set it up to make it work.