5 ms·
The benchmark you linked was to "programming performance", not generic LLM "intelligence". The situation for the little guy is wildly better than most people i
by PostOnce 3y ago
The benchmark you linked was to "programming performance", not generic LLM "intelligence".
The situation for the little guy is wildly better than most people imagine.
- moffkalast 3y agoYep, that's what I'm saying, programming performance is seemingly very indicative of model inteligence (assuming it's tuned well enough to be able to run the benchmark at all). Coding is an exercise in problem solving and abstract thinking after all. There are exceptions of course, as there are a few models (e.g. Vicuna, Baize) that don't do well at coding at all but otherwise perform well for chat, and the coding models I mentioned that game the benchmark by sacrificing performance in all other areas. If you exclude those, it's very a accurate overall reasoning level comparison, at least it fits most to what I've seen their performance was for various tasks when testing out individual models. The only other valid benchmark that isn't coding are the SAT and LSAT tests that OpenAI runs on all of their models, but afaik there isn't an open version that would be widely used.