3 ms·
I really don't like the snarky tone of the parent comment. Nonetheless, I don't think this is even something that can easily be benchmarked. I'd recommend you
by PartiallyTyped 2y ago
I really don't like the snarky tone of the parent comment.
Nonetheless, I don't think this is even something that can easily be benchmarked. I'd recommend you take a look at aider [1], and consider how I drew similarities between it and what's presented here.
Has ClosedAI presented any benchmarks / evaluation protocols?
[1] https://aider.chat/ https://aider.chat/
- tmnvdb 2y agoYes, they show benchmarks in the article linked here. Did you not read it?
- PartiallyTyped 2y agoI don’t think you actually read it. The benchmarks are in reference to the model that’s underlying deep-research, and not deep-research itself. For the latter, they have anecdata from scientists.