3 ms·
The majority of these LLMs are not cutting edge, and many of them were designed for specific purposes other than answering prompts like these. I won't defend th
by caturopath 3y ago
The majority of these LLMs are not cutting edge, and many of them were designed for specific purposes other than answering prompts like these. I won't defend the level of hype coming from many corners, but it isn't fair to look at these responses to get the ceiling on what LLMs can do -- for that you want to look at only the best (GPT4, which is represented, and Bard, which isn't, essentially). Claude 2 (also represented) is in the next tier. None of the other models are at their level, yet.
You'd also want to look at models that are well-suited to what you're doing -- some of these are geared to specific purposes. Folks are pursuing the possibility that the best model would fully-internally access various skills, but it isn't known whether that is going to be the best approach yet. If it isn't, selecting among 90 (or 9 or 900) specialized models is going to be a very feasible engineering task.
> The 12-bar blues progressions seem mostly clueless.
I mean, it's pretty amazing that they many look coherent compared to the last 60 years of work at making a computer talk to you.
That being said, I played GPT4's chords and they didn't sound terrible. I don't know if they were super bluesy, but they weren't _not_ bluesy. If the goal was to build a music composition assistant tool, we can certainly do a lot better than any of these general models can do today.
> The question is will any of these ever get significantly better with time, or are they mostly going to stagnate?
No one knows yet. Some people think that GPT4 and Bard have reached the limits of what our datasets can get us, some people think we'll keep going on the current basic paradigm to AGI superintelligence. The nature of doing something beyond the limits of human knowledge, creating new things, is that no one can tell you for sure the result.
If they do stagnate, there are less sexy ways to make models perform well for the tasks we want them for. Even if the models fundamentally stagnate, we aren't stuck with the quality of answers we can get today.