4 ms·
> This is not scientific at all, just vibes, YMMV. This is the problem. I would love to have a product sheet showing what each models strengths an weaknesses
by dkersten 4mo ago
> This is not scientific at all, just vibes, YMMV.
This is the problem.
I would love to have a product sheet showing what each models strengths an weaknesses are, so that I can have a clear decision tree of "if this kind of work, use model X", or "model Y should be used in ways Z". But they all look the same from the outside and the only way to figure out which might be marginally better at what is to do extensive, time consuming, and perhaps expensive testing.
- amelius 4mo agoYes, but benchmarks can be gamed. Maybe we need better reviewers then?
- dotancohen 4mo agoHonestly, the differences between AI models always felt to me like the differences between coworkers or job candidates. They don't all share the same strengths and weaknesses - and they all have both good days and bad days. Realising this made me respect the "I" in "AI" a bit more seriously.
- couscouspie 4mo agoThat would be ideal, but AI is less like a tool and more like a human in this regard and you don't have character sheets for each of your colleagues, as well.
- bluegatty 4mo agoThese are $1 Trillion dollar companies that can't produce explicit details on how their products work? It's nonsense.
- deleted 4mo ago[deleted]
- sixothree 4mo agoI think if they could explain how they work, their strengths and weaknesses, they would reveal to the world whose data they've been appropriating.
- bluegatty 4mo agoThat's another thing altogether. They can characterize the behaviour without quite giving up who and where the data comes from. Admittedly, yes, there's some overlap there. They would have to admit 'seen it in the training data' as a factor, and that opens a can of worms.
- supergarfield 4mo agoIf my coworker was part of a clone series of 100 million units, requesting a character sheet would be pretty reasonable
- coldtea 4mo ago>I would love to have a product sheet showing what each models strengths an weaknesses are, so that I can have a clear decision tree of "if this kind of work, use model X", or "model Y should be used in ways Z". But they all look the same from the outside and the only way to figure out which might be marginally better at what is to do extensive, time consuming, and perhaps expensive testing. Think of it less like a static tool, and more like a human helper, where the same holds.
- madeofpalk 4mo agoExcept, where every different model and version is like a different person where you need to learn their idiosyncrasies of how they work every other month. It's a very very bizarre way to use a computer. Personally, I just don't. I'll use and prompt the LLMs the way that feels natural to me and move on with my life. Maybe I don't always get completely optimal results from them, but im also not spending half my day pleading with the computer to do a task.
- user43928 4mo agoI also don't think I need to prompt Claude differently than Codex. The most important thing to be aware of in my opinion would be that Claude is better at UI design, and leaves a lot more comments in the code. Other than that the results seem similar, at least functionally. I do not usually review the code style.
- ACCount37 4mo agoOne issue with that is that human helpers last longer. LLMs cycle in and out in months, and what held for Your Favorite LLM 6.7 may not hold for Your Favorite LLM 6.9.
- renegade-otter 4mo agoRight, this is why I would slam the breaks on investing into your workflow all of your time and effort, because 2 months from now it may be out the window. Frontier models are also constantly being tweaked, so what worked yesterday may be off today. ChatGPT was obedient with the grill-me technique, just wrote a plan. Yesterday it started jumping to implementation. Why?
- epolanski 4mo agoThe problem is that this is very hard to replicate and benchmarks focus on E2E tests, going from one prompt to the final solution. They do not test how models perform when used interactively, like most of us do.
- yunohn 4mo ago> a product sheet showing what each models strengths an weaknesses are This presumes that the labs themselves know how well their models perform. But all they have are overtuned benchmarks and hype vibes.
- egwor 4mo agoMaybe this is similar to web search too. We know how to get google to return the results we want, and when we use other tools like Bing we get other behaviour.
- m-dot-reviews 4mo agoSo, this may not be precisely what you're looking for but it may come close. I've put together a simple site for sharing ratings/opinions on models on a task-specific granularity. https://model.reviews/ https://model.reviews/ The idea is that benchmark score comparisons are useful for a large cross-product comparison across models + their settings, but less useful if you're looking for the best model for <your-specific-task>. So on this site, each model gets its own page showing the list of tasks that people have rated it on, and the score out of 10 for each task. Common tasks, like coding, will likely be on most/all models, and more niche tasks may only be on a few. It is human moderated (by me only right now). The corpus is pretty empty right now, so please spread the word if this seems like a useful idea!
- demades 4mo ago[flagged]