2 ms·
Not yet, but starting the runs for 3.7 later today! The cost for running all the evals (across all models) was about $10k. Simply giving the agents access to t
by noddybear 2y ago
Not yet, but starting the runs for 3.7 later today! The cost for running all the evals (across all models) was about $10k.
Simply giving the agents access to the tool descriptions and API schema is like 20k tokens from the outset.
It would be really cool to use retrieval techniques to reduce this burden. I suspect that this will also outright improve the performance of all models - which becomes worse as the context scales.
- 0xmason 2y agoPlease add the results of the eval to your leaderboard / github! Looking forward to it. GPT 4.5 seems prohibitively expensive here though lol.
- stared 2y agoSince Claude 3.5 Sonnet is that good, I am curious how fares Claude 3.5 Haiku. For programming-like tasks, I expect similar-ish distribution that in programming, see e.g. https://web.lmarena.ai/leaderboard https://web.lmarena.ai/leaderboard