2 ms·
Hi, I am the author, I completely agree! I set out to run a vibe test on this one, not a benchmark, the real benchmarks are listed. My test shows what the model
by jameswhitford 3mo ago
Hi, I am the author, I completely agree! I set out to run a vibe test on this one, not a benchmark, the real benchmarks are listed. My test shows what the models can do when both tasked with a long-running, technically difficult, one-shot task.
I think your test you describe (collaborative, task delegation, task completion, TTD, steerability) is a great format for a future test that I will definitely try out.
- meander_water 3mo agoThanks, I didn't mean to be brusque, but I have seen a lot of these vibe tests lately that come to grand conclusions like "X model is better than Y" from the result of a single prompt. Appreciate you sharing the results of your tests though!
- jameswhitford 3mo agoI appreciate the feedback!
- wongarsu 3mo agoTbf, most of the "real benchmarks" have issues that are just as bad. Assessing LLM performance is just hard
- oceansky 3mo agoAnd personal too. Different engineers are using them for different use cases.
- ramraj07 3mo agoThe important point is that your benchmark is pretty much irrelevant for the actual usage. Thus whatever conclusion you draw is not just irrelevant but misleading.